arrow
返回

Diverse Temporal Aggregation and Depthwise Spatiotemporal Factorization for Efficient Video Classification

delete2021-01-01
delete2
delete
OA
AI
Y
Youngwan Lee
H
Hyung-Il Kim
K
Kimin Yun
J
Jinyoung Moon *
DOI:10.1109/ACCESS.2021.3132916delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Video classification researches have recently attracted attention in the fields of temporal modeling and efficient 3D convolutional architectures. However, the temporal modeling methods are not efficient, and there is little interest in how to deal with temporal modeling in the 3D efficient architectures. To build an efficient 3D architecture for temporal modeling, we propose a new 3D backbone network, called VoV3D, that consists of a temporal one-shot aggregation (T-OSA) module and a depthwise factorized component, D(2 + 1)D. The T-OSA is devised to build a feature hierarchy by aggregating spatiotemporal features with different temporal receptive fields. Stacking this T-OSA enables the network itself to model short-range as well as long-range temporal relationships across frames without any external modules. We also design a depthwise spatiotemporal factorization module, D(2 + 1)D, that decomposes a 3D depthwise convolution into two spatial and temporal depthwise convolutions for efficient architecture. Through the proposed temporal modeling method (T-OSA) and the efficient factorization module (D(2 + 1)D), we construct two types of VoV3D networks: VoV3D-M and VoV3D-L. Thanks to its efficiency and effectiveness of their temporal modeling, VoV3D-L has 4 fi fewer model parameters and 14 fi less computation, surpassing the state-of-the-art TEA model on both Something-Something and Kinetics-400 datasets. We hope that VoV3D can serve as a baseline for efficient temporal modeling architecture.
Keyword:
Action recognition
video classification
temporal modeling
efficient 3D CNN architecture
spatial-temporal feature

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

引用论文

引用论文

err
IF0
err
err0
PREAI
err
err分享
err收藏
Time-dependent depolarization of aligned HD molecules
err2009-01-01
err0
errOAAI
errNate C.-M. Bartlett; Daniel J. Miller; Richard N. Zare; Andrew J. Alexander; Dimitris Sofikitis; T. Peter Rakitzis
err分享
err收藏
The mental health of staff working in intensive care during COVID-19
err
IF0
err2020-11-04
err0
errOAAI
errNeil Greenberg; Dale Weston; Charlotte Hall; Tristan Caulfield; Victoria Williamson; Kevin Fong
err分享
err收藏
ImageNet Large Scale Visual Recognition ChallengeImageNet大规模视觉识别挑战
err2015-04-11
err2.7W
PREAI
errRussakovsky, Olga; Deng, Jia; Su, Hao; Krause, Jonathan; Satheesh, Sanjeev; Ma, Sean; Huang, Zhiheng; Karpathy, Andrej; Khosla, Aditya; Bernstein, Michael; Berg, Alexander C.; Fei-Fei, Li
err分享
err收藏
学者 查看更多内容