arrow
返回

A component-based video content representation for action recognition

delete2019-10-01
delete11
PRE
AI
V
Vida Adeli
E
Ehsan Fazl-Ersi *
A
Ahad Harati
DOI:10.1016/j.imavis.2019.08.009delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
This paper investigates the challenging problem of action recognition in videos and proposes a new component-based approach for video content representation. Although satisfactory performance for action recognition has already been obtained for certain scenarios, many of the existing solutions require fully-annotated video datasets in which region of the activity in each frame is specified by a bounding box. Another group of methods require auxiliary techniques to extract human-related areas in the video frames before being able to accurately recognize actions. In this paper, a Weakly-Supervised Learning (WSL) framework is introduced that eliminates the need for per-frame annotations and learns video representations that improve recognition accuracy and also highlights the activity related regions within each frame. To this end, two new representation ideas are proposed, one focus on representing the main components of an action, i.e. actionness regions, and the other focus on encoding the background context to represent general and holistic cues. A three-stream CNN is developed, which takes the two proposed representations and combines them with a motion-encoding stream. Temporal cues in each of the three different streams are modeled through LSTM, and finally fully-connected neural network layers are used to fuse various streams and produce the final video representation. Experimental results on four challenging datasets, demonstrate that the proposed Component-based Multi-stream CNN model (CM-CNN), trained on a WSL setting, outperforms the state-of-the-art in action recognition, even the fully-supervised approaches. (C) 2019 Elsevier B.V. All rights reserved.
Keyword:
Actionness likelihood
Action recognition
Action components
LSTM
Three-stream convolutional neural network
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Image and Vision Computing 封面图
Image and Vision Computing
IF:
4.2
论文数:
4.1K
被引数:
6.7K

机构

F
Ferdowsi University Mashhad
学者数:
8.0K
论文数: 7.4K
被引数: 44
引用论文

引用论文

Human Action Recognition Using Adaptive Local Motion Descriptor in Spark
err2017-01-01
err31
errOAAI
errUddin, M. D. Azher; Joolee, Joolekha Bibi; Alam, Aftab; Lee, Young-Koo
err分享
err收藏
Calibration of binocular vision measurement system
err2016-01-01
err0
PREAI
err杨景豪 YANG Jing-hao; 刘 巍 LIU Wei; 刘 阳 LIU Yang; 王福吉 WANG Fu-ji; 贾振元 JIA Zhen-yuan
err分享
err收藏
Action detection based on tracklets with the two-stream CNN
err2017-08-24
err16
PREAI
errZhang, Minwen; Gao, Chenqiang; Li, Qiang; Wang, Lan; Zhang, Jiayao
err分享
err收藏
Multi-stream CNN: Learning representations based on human-related regions for action recognition
err2018-07-01
err184
errOAAI
errTu, Zhigang; Xie, Wei; Qin, Qianqing; Poppe, Ronald; Veltkamp, Remco C.; Li, Baoxin; Yuan, Junsong
err分享
err收藏
err分享
err收藏
没有更多内容