Return
Information-Bottleneck Guided Hybrid Neural Architecture Search for Temporal Action Detection in Untrimmed Videos
Y
M
L
S
DOI:10.1109/tip.2026.3720512.png)
Abstract
En 中文
Temporal Action Detection (TAD) in untrimmed videos requires effective spatial feature extraction for precise action classification and temporal feature modeling for accurate boundary localization. To achieve effective spatio-temporal feature integration, several works manually design rule-based (i.e., sequential or parallel) hybrid Mamba-Transformer networks for TAD. However, few studies explore diverse integration strategies and network topologies due to the inherent limitations of manual design. Therefore, we propose NAS-TAD, the first Neural Architecture Search framework for TAD, systematically exploring this untouched problem. Specifically, we develop a spatio-temporal NAS objective function based on information-bottleneck theory to quantify task-relevant spatio-temporal features, providing interpretable guidance for the network search and optimization process. Furthermore, we reformulate Transformer self-attention as a state-space model, thereby enabling seamless switching between Mamba and Transformer blocks in a unified weight-sharing search space. Consequently, comprehensive experiments on ActivityNet, THUMOS14, HACS and FineAction demonstrate the effectiveness of the searched hybrid architectures, providing new insights into temporal and spatial feature fusion for TAD. Code is available for reproduction at https://anonymous.4open.science/r/nastad.
Keywords:
Temporal action detection
Neural architecture search
Information bottleneck theory
Video understanding
Journal
IF:
13.7
Papers:
1.0W
Citations:
8.4W
