返回
Spatio-Temporal Collaborative Module for Efficient Action Recognition
DOI:10.1109/TIP.2022.3221292.png)
摘要
En 中文
Efficient action recognition aims to classify a video clip into a specific action category with a low computational cost. It is challenging since the integrated spatial-temporal calculation (e.g., 3D convolution) introduces intensive operations and increases complexity. This paper explores the feasibility of the integration of channel splitting and filter decoupling for efficient architecture design and feature refinement by proposing a novel spatio-temporal collaborative (STC) module. STC splits the video feature channels into two groups and separately learns spatio-temporal representations in parallel with decoupled convolutional operators. Particularly, STC consists of two computation-efficient blocks, i.e., ST and Ts, where they extract either spatial (S.) or temporal (T.) features and further refine their features with either temporal (.(T)) or spatial (.(S)) contexts globally. The spatial/temporal context refers to information dynamics aggregated from temporal/spatial axis. To thoroughly examine our method's performance in video action recognition tasks, we conduct extensive experiments using five video benchmark datasets requiring temporal reasoning. Experimental results show that the proposed STC networks achieve a competitive trade-off between model efficiency and effectiveness.
Keyword:
Efficient action recognition
deep video neural network
channel split
feature contextualization
期刊
IF:
13.7
论文数:
1.0W
被引数:
8.4W
机构
引用论文
A silicotungstate-based copper–viologen hybrid photocatalytic compound for efficient degradation of organic dyes under visible light
CrystEngComm
IF0
Synthesis of n-type semiconducting diamond film using diphosphorus pentaoxide as the doping source以五氧化二磷为掺杂源合成n型半导体金刚石膜

