返回
Three-stream CNNs for action recognition
DOI:10.1016/j.patrec.2017.04.004.png)
摘要
En 中文
Existing Convolutional Neural Networks (CNNs) based methods for action recognition are either spatial or temporally local while actions are 3D signals. In this paper, we propose a global spatial-temporal three-stream CNNs architecture, which is able to be used for action feature extraction. Specifically, the three-stream CNNs comprises of spatial, local temporal and global temporal streams generated respectively from deep learning single frame, optical flow and global accumulated motion features in the form of a new formulation named Motion Stacked Difference Image (MSDI). Moreover, a novel soft Vector of Locally Aggregated Descriptors (soft-VLAD) is developed to further represent the extracted features, combining the advantage of Gaussian Mixture Models (GMMs) and VLAD by encoding data according to their overall probability distribution and the corresponding difference with respect to clustered centers. To deal with the inadequacy of training samples during learning, we introduce a data augmentation scheme which is very efficient due to its origin. at cropping across videos. We conduct our experiments on UCF101 and HMDB51 datasets, and the results demonstrate the effectiveness of our approach. (C) 2017 Elsevier B.V. All rights reserved.
Keyword:
Action recognition
Three-stream convolutional neural networks
Soft Vector of Locally Aggregated
Descriptors
Support Vector Machines
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.3
论文数:
7.9K
被引数:
1.6W
机构
引用论文
ARCH: Adaptive recurrent-convolutional hybrid networks for long-term action recognition
NEUROCOMPUTING
IF6.5

