arrow
返回

MAR: Masked Autoencoders for Efficient Action Recognition

delete2024-01-01
delete14
delete
OA
AI
Z
Zhiwu Qing
S
Shiwei Zhang *
Z
Ziyuan Huang
X
Xiang Wang
王岳环 (Yuehuan Wang)
Y
Yiliang Lv
C
Changxin Gao
N
Nong Sang *
DOI:10.1109/TMM.2023.3263288delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Standard approaches for video action recognition usually operate on full input videos, which is inefficient due to the widespread spatio-temporal redundancy in videos. The recent progress in masked video modelling, specifically VideoMAE, has shown the ability of vanilla Vision Transformers (ViT) to complement spatio-temporal contexts using limited visual content. Inspired by this, we propose Masked Action Recognition (MAR), which reduces redundant computation by discarding a proportion of patches and operating only on a portion of the videos. MAR includes two essential components: cell running masking and bridging classifier. Specifically, to enable the ViT to perceive the details beyond the visible patches, cell running masking is used to preserve the spatio-temporal correlations in videos. This ensures that the patches at the same spatial location can be observed in turn for easy reconstructions. Additionally, we notice that, although the partially observed features can reconstruct semantically explicit invisible patches, they fail to achieve accurate classification. To address this issue, we propose a bridging classifier that can help fill the semantic gap between the ViT encoded features used for reconstruction and the specialized features used for classification. Our proposed MAR can reduce the computational cost of ViT by 53%. Extensive experiments have demonstrated that MAR consistently outperforms existing ViT models by a notable margin. Notably, we found that a ViT-Large model fine-tuned by MAR achieves comparable performance to a ViT-Huge model fine-tuned by standard training methods on both Kinetics-400 and Something-Something v2 datasets. Moreover, the computation overhead of our ViT-Large model is only 14.5% of that of the ViT-Huge model.
Keyword:
Efficient action recognition
masked autoencoders
spatio-temporal redundancy
vision transformer

期刊

IEEE Transactions on Multimedia 封面图
IEEE Transactions on Multimedia
IF:
9.7
论文数:
4.5K
被引数:
2.4W

机构

A
alibaba group
学者数:
1.1K
论文数: 789
被引数: 0
N
National University of Singapore
学者数:
7.6W
论文数: 6.5W
被引数: 11.4W
引用论文

引用论文

Associations between electroencephalographic and magnetic resonance imaging findings in tuberous sclerosis complex
err2009-12-01
err0
PREAI
errAnne Gallagher; Catherine J. Chu-Shore; Maria A. Montenegro; Philippe Major; Daniel J. Costello; David A. Lyczkowski; David Muzykewicz; Colin Doherty; Elizabeth A. Thiele
err分享
err收藏
err分享
err收藏
p8 Deficiency Causes Siderosis in Spleens and Lymphocyte Apoptosis in Acute Pancreatitis
err2014-11-01
err0
PREAI
errSebastian Weis; Tilmann Cornelius Schlaich; Faramarz Dehghani; Tânia Carvalho; Ines Sommerer; Stephan Fricke; Franka Kahlenberg; Joachim Mössner; Albrecht Hoffmeister
err分享
err收藏
Time-dependent depolarization of aligned HD molecules
err2009-01-01
err0
errOAAI
errNate C.-M. Bartlett; Daniel J. Miller; Richard N. Zare; Andrew J. Alexander; Dimitris Sofikitis; T. Peter Rakitzis
err分享
err收藏
err分享
err收藏
ParamCrop: Parametric Cubic Cropping for Video Contrastive Learning
err2023-01-01
err2
PREAI
errQing, Zhiwu; Huang, Ziyuan; Zhang, Shiwei; Tang, Mingqian; Gao, Changxin; Jin, Rong; Ang Jr, Marcelo H.; Sang, Nong
err分享
err收藏
学者 查看更多内容