返回
Action Recognition with Temporal Scale-Invariant Deep Learning Framework
DOI:10.1109/CC.2017.7868164.png)
摘要
En 中文
Recognizing actions according to video features is an important problem in a wide scope of applications. In this paper, we propose a temporal scale-invariant deep learning framework for action recognition, which is robust to the change of action speed. Specifically, a video is firstly split into several sub-action clips and a keyframe is selected from each sub-action clip. The spatial and motion features of the keyframe are extracted separately by two Convolutional Neural Networks (CNN) and combined in the convolutional fusion layer for learning the relationship between the features. Then, Long Short Term Memory (LSTM) networks are applied to the fused features to formulate long-term temporal clues. Finally, the action prediction scores of the LSTM network are combined by linear weighted summation. Extensive experiments are conducted on two popular and challenging benchmarks, namely, the UCF-101 and the HMDB51 Human Actions. On both benchmarks, our framework achieves superior results over the state-of-the-art methods by 93.7% on UCF-101 and 69.5% on HMDB51, respectively.
Keyword:
action recognition
CNN
LSTM
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.1
论文数:
1.9K
被引数:
5.0K
机构
引用论文
暂无论文信息


