arrow
返回

A Generic Framework for Video Annotation via Semi-Supervised Learning

delete2012-08-01
delete49
PRE
AI
张
张天柱 (Tianzhu Zhang) *
徐
徐常胜 (Changsheng Xu)
朱光玉 封面图
朱光玉 (Guangyu Zhu)
S
Si Liu
卢
卢汉清 (Hanqing Lu)
DOI:10.1109/TMM.2012.2191944delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Learning-based video annotation is essential for video analysis and understanding, and many various approaches have been proposed to avoid the intensive labor costs of purely manual annotation. However, there lacks a generic framework due to several difficulties, such as dependence of domain knowledge, insufficiency of training data, no precise localization and inefficacy for large-scale video dataset. In this paper, we propose a novel approach based on semi-supervised learning by means of information from the Internet for interesting event annotation in videos. Concretely, a Fast Graph-based Semi-Supervised Multiple Instance Learning (FGSSMIL) algorithm, which aims to simultaneously tackle these difficulties in a generic framework for various video domains (e. g., sports, news, and movies), is proposed to jointly explore small-scale expert labeled videos and large-scale unlabeled videos to train the models. The expert labeled videos are obtained from the analysis and alignment of well-structured video related text (e. g., movie scripts, web-casting text, close caption). The unlabeled data are obtained by querying related events from the video search engine (e. g., YouTube, Google) in order to give more distributive information for event modeling. Two critical issues of FGSSMIL are: 1) how to calculate the weight assignment for a graph construction, where the weight of an edge specifies the similarity between two data points. To tackle this problem, we propose a novel Multiple Instance Learning Induced Similarity (MILIS) measure by learning instance sensitive classifiers; 2) how to solve the algorithm efficiently for large-scale dataset through an optimization approach. To address this issue, Concave-Convex Procedure (CCCP) and nonnegative multiplicative updating rule are adopted. We perform the extensive experiments in three popular video domains: movies, sports, and news. The results compared with the state-of-the-arts are promising and demonstrate the effectiveness and efficiency of our proposed approach.
Keyword:
Broadcast video
concave-convex procedure (CCCP)
event detection
graph
Internet
multiple instance learning
semi-supervised learning
web-casting text
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Multimedia 封面图
IEEE Transactions on Multimedia
IF:
9.7
论文数:
4.5K
被引数:
2.4W

机构

I
institute of automation, cas
学者数:
2.2K
论文数: 2.1K
被引数: 2
C
chinese academy of sciences
学者数:
56.7W
论文数: 45.0W
被引数: 704
引用论文

引用论文

Mesoanalysis of the Big Thompson Storm
err1979-01-01
err0
errOAAI
errFernando Caracena; Robert A. Maddox; L. Ray Hoxit; Charles F. Chappell
err分享
err收藏
err分享
err收藏
Structural Isomerization of the Gas‐Phase 2‐Norbornyl Cation Revealed with Infrared Spectroscopy and Computational Chemistry
err2014-05-07
err0
PREAI
errJonathan D. Mosley; Justin W. Young; Jay Agarwal; Henry F. Schaefer; Paul v. R. Schleyer; Michael A. Duncan
err分享
err收藏
Lactic acidosis
err1986-03-01
err0
errOAAI
errNicolaos E. Madias
err分享
err收藏
err分享
err收藏
学者 查看更多内容