arrow
返回

Exploiting Mid-Level Semantics for Large-Scale Complex Video Classification

delete2019-10-01
delete15
PRE
AI
J
Ji Zhang
梅魁志 (Kuizhi Mei) *
Y
Yu Zheng
J
Jianping Fan
DOI:10.1109/TMM.2019.2907453delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
As the amount of available video data has grown substantially, automatic video classification has become an urgent yet challenging task. Most video classification methods focus on acquiring discriminative spacial visual features and motion patterns for video representation, especially deep learning methods, which have achieved very good results on action recognition problems. However, the performance of most of these methods drastically degenerates for more generic video classification tasks where the video contents are much more complex. Thus, in this paper, the mid-level semantics of videos are exploited to bridge the semantic gap between low-level features and high-level video semantics. Inspired by the term frequency-inverse document frequency, a word weighting method for the problem of text classification is introduced to the video domain. The visual objects in videos are regarded as the words in texts, and two new weighting methods are proposed to encode videos by weighting visual objects according to the characteristics of videos. In addition, the semantic similarities between video categories and visual objects are introduced from the text domain as privileged information to facilitate classifier training on the obtained semantic representations of videos. The proposed semantic encoding method (semantic stream) is then fused with the popular two-stream CNN model for the final classification results. Experiments are conducted on two large-scale complex video datasets, CCV and ActivityNet. The experimental results validate the effectiveness of the proposed methods.
Keyword:
Mid-level semantics
word weighting methods
semantic similarities
learning using privileged information
large-scale video classification
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Multimedia 封面图
IEEE Transactions on Multimedia
IF:
9.7
论文数:
4.5K
被引数:
2.4W

机构

X
xi'an jiaotong university
学者数:
9.3W
论文数: 6.7W
被引数: 75
U
university of north carolina
学者数:
7.4W
论文数: 6.5W
被引数: 93
X
Xidian University
学者数:
2.4W
论文数: 1.9W
被引数: 9.7K
学者 查看更多机构