Return
A renaissance of explicit motion information mining from transformers for action recognition
DOI:10.1016/j.patcog.2025.112645.png)
Abstract
En 中文
• We comprehensively compare the differences between cost volume and self-attention. Furthermore, we summarize and empirically validate three key properties of cost volume in effective motion modeling. • We propose the Explicit Motion Information Mining module (EMIM) to enrich the motion modeling capacity for the existing transformer. The EMIM module preserves the original capacity for contextual aggregation and develops a new ability in effective motion modeling. • We validate the effectiveness of the proposed method on several widelyused datasets, and achieve state-of-the-art performances, especially on motion-sensitive datasets.
Journal
IF:
7.6
Papers:
1.3W
Citations:
4.5W

