arrow
Return

Rethinking Lightweight: Multiple Angle Strategy for Efficient Video Action Recognition

delete2022-01-01
delete3
PRE
AI
J
Jianyu Chen
王中元 (Zhongyuan Wang)
K
Kangli Zeng
Z
Zheng He *
Z
Zixiang Xiong
DOI:10.1109/LSP.2022.3144074delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Video action recognition task involves modeling spatiotemporal information, and efficiency is critical to capture spatiotemporal dependencies in the video. Most existing models rely on optical flow information to capture the dynamic visual tempos between consecutive video frames. Although impressive performance can be achieved by combining optical flow with RGB, the time-consuming nature of optical flow computation cannot be ignored. Moreover, 3D CNN has successfully modeled spatiotemporal information, yet the enormous computational volume is unsuitable for real-time action recognition. In this letter, we propose a novel lightweight video feature extraction strategy that achieves better recognition performance with lower FLOPs. In particular, we perform convolution on the video cube from three orthogonal angles to learn its appearance and motion features. Compared with the computational volume of 3D CNN, our proposed method is more economical and thus meets the lightweight requirements. Extensive experimental results on public Something Something-V1&V2 and Diving48 datasets show our approach achieves the state-of-the-art performance.
Keywords:
Spatiotemporal phenomena
Convolution
Feature extraction
Three-dimensional displays
Computational modeling
Solid modeling
Biological system modeling
Multiple dimensions
separated convolution
spatiotemporal information
video action recognition

Journal

IEEE Signal Processing Magazine cover
IEEE Signal Processing Magazine
IF:
9.6
Papers:
1.1W
Citations:
1.7W

Organization

T
Texas A&M University System
Scholars:
4.4W
Papers: 4.0W
Citations: 4.0K
W
wuhan university
Scholars:
8.1W
Papers: 5.8W
Citations: 70