arrow
返回

Efficient Video Transformers via Spatial-temporal Token Merging for Action Recognition

delete2024-01-11
delete0
PRE
AI
Z
Zhanzhou Feng *
J
Jiaming Xu
马雷 封面图
马雷 (Лей Ма)
张史梁 封面图
张史梁 (Shiliang Zhang)
DOI:10.1145/3633781delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Transformer has exhibited promising performance in various video recognition tasks but brings a huge computational cost in modeling spatial-temporal cues. This work aims to boost the efficiency of existing video transformers for action recognition through eliminating redundancies in their tokens and efficiently learning motion cues of moving objects. We propose a lightweight and plug-and-play module, namely Spatial-temporal Token Merger (STTM), to merge the tokens belonging to the same object into a more compact representation. STTM first adaptively identifies crucial object clues underlying the video as meta tokens. Similarity scores between input tokens and meta tokens are hence computed and used to guide the fusion of similar tokens in both spatial and temporal domains, respectively. To compensate for motion cues lost in the merging procedure, we compute the linear aggregation of spatial-temporal positions of tokens as motion features. STTM hence outputs a compact set of tokens fusing both appearance and motion features of moving objects. This procedure substantially decreases the number of tokens that need to be processed by each Transformer block and boosts the efficiency. As a general module, STTM can be applied to different layers of various video Transformers. Extensive experiments on the action recognition datasets Kinectics-400 and SSv2 demonstrate its promising performance. For example, it reduces the computation complexity of ViT by 38% while maintaining a similar performance on Kinectics-400. It also brings 1.7% gains of top-1 accuracy on SSv2 under the same computational cost.
Keyword:
Efficient video recognition
deep learning
transformer
spatial-temporal information

期刊

ACM Transactions on Multimedia Computing Communications and Applications 封面图
ACM Transactions on Multimedia Computing Communications and Applications
IF:
6
论文数:
2.0K
被引数:
5.4K

机构

P
peking university
学者数:
11.9W
论文数: 8.7W
被引数: 146
引用论文

引用论文

Time-dependent depolarization of aligned HD molecules
err2009-01-01
err0
errOAAI
errNate C.-M. Bartlett; Daniel J. Miller; Richard N. Zare; Andrew J. Alexander; Dimitris Sofikitis; T. Peter Rakitzis
err分享
err收藏
The mental health of staff working in intensive care during COVID-19
err
IF0
err2020-11-04
err0
errOAAI
errNeil Greenberg; Dale Weston; Charlotte Hall; Tristan Caulfield; Victoria Williamson; Kevin Fong
err分享
err收藏
Using SEPIC Topology for Improving Power Factor in Distributed Power Supply Systems
err2015-09-22
err0
PREAI
errJ. Sebastián; J. Uceda; J.A. Cobos; J. Arau
err分享
err收藏
The effect of autonomic arousal on attentional focus
err2000-12-01
err0
PREAI
errJ I. Tracy; F Mohamed; S Faro; R Tiver; A Pinus; C Bloomer; A Pyrros; J Harvan
err分享
err收藏
err分享
err收藏
err分享
err收藏
Viologen‐Based Uranyl Coordination Polymers: Anion‐Induced Structural Diversity and the Potential as a Fluorescent Probe
err2021-12-09
err0
PREAI
errKong‐Qiu Hu; Li‐Wen Zeng; Xiang‐He Kong; Zhi‐Wei Huang; Ji‐Pan Yu; Lei Mei; Zhi‐Fang Chai; Wei‐Qun Shi
err分享
err收藏
学者 查看更多内容