arrow
Return

A renaissance of explicit motion information mining from transformers for action recognition

delete2025-10-31
delete0
PRE
AI
P
Peiqin Zhuang
白磊(LeiBai) (Lei Bai)
Y
Yichao Wu
D
Ding Liang
L
Luping Zhou
Y
Yali Wang
W
Wanli Ouyang
DOI:10.1016/j.patcog.2025.112645delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• We comprehensively compare the differences between cost volume and self-attention. Furthermore, we summarize and empirically validate three key properties of cost volume in effective motion modeling. • We propose the Explicit Motion Information Mining module (EMIM) to enrich the motion modeling capacity for the existing transformer. The EMIM module preserves the original capacity for contextual aggregation and develops a new ability in effective motion modeling. • We validate the effectiveness of the proposed method on several widelyused datasets, and achieve state-of-the-art performances, especially on motion-sensitive datasets.

Journal

Pattern Recognition cover
Pattern Recognition
IF:
7.6
Papers:
1.3W
Citations:
4.5W

Organization

S
shanghai ai laboratory
Scholars:
48
Papers: 30
Citations: 0
S
sensetime group ltd.
Scholars:
2
Papers: 1
Citations: 0
T
The University of Sydney
Scholars:
2.6K
Papers: 1.1K
Citations: 7.9W
S
Shenzhen Institutes of Advanced Technology
Scholars:
821
Papers: 286
Citations: 0
researcher View more organizations