1
Return

Towards Robust Temporal Action Detection: Benchmark and A Strong Baseline

delete2026-08-01
delete0
PRE
AI
R
Runhao Zeng *
J
Jiaming Liang
J
Jiaqi Mao
X
Xiaoyong Chen
王维 (Wei Wang)
Y
Yong Guo
L
Limin Wang
V
Victor C. M. Leung
X
Xiping Hu
DOI:10.1007/s11263-026-02956-3delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Temporal action detection (TAD) aims to locate action positions and recognize action categories in long-term untrimmed videos. Despite the fact that numerous methods have achieved encouraging outcomes, their robustness has not been comprehensively investigated. In practice, it is observed that temporal information in videos can sometimes be compromised, such as by missing or blurred frames. Notably, existing methods are highly vulnerable to these scenarios, often experiencing a significant decline in performance even when only a single frame is corrupted. In this paper, we take the first step towards benchmarking the temporal robustness of TAD models and aim to identify the factors influencing temporal robustness. This, in turn, provides insights for designing more robust TAD models. To formally assess robustness, we establish three temporal corruption robustness benchmarks, namely THUMOS14-C, ActivityNet-1.3-C and MultiTHUMOS-C, which consider eight common types of corruption encountered during video recording and transmission. Each type of corruption is applied at three levels of severity, resulting in a total of 24 distinct corruptions, which comprehensively cover different durations of corruption, serving as a controllable representative of practical real-world scenarios. On these benchmarks, we conduct an extensive analysis of the robustness of 12 leading TAD methods and uncover several noteworthy findings: 1) Existing methods are particularly vulnerable to temporal corruptions, with end-to-end methods likely being more susceptible than those employing a pre-trained feature extractor on THUMOS14-C; 2) The primary source of vulnerability is localization error rather than classification error; and 3) TAD models tend to exhibit the most significant performance degradation when corruptions occur in the middle of an action instance. Furthermore, we investigate the impact of diverse TAD model designs on temporal robustness by evaluating eight key factors across three crucial stages: feature representation, architecture design, and training strategy. Based on the insights gained from these explorations, we propose a recipe with six steps to build a strong TAD baseline. Experiments conducted on three benchmark datasets demonstrate that our recipe not only improves robustness against corruption but also results in enhancements on clean data. Specifically, on the THUMOS14-C dataset, we achieve an 11.56% improvement in relative robustness and a 1.32% increase in clean mAP. We believe that this study will play a crucial role in shaping future research on robust video analysis. The benchmark dataset is available at https://github.com/Alvin-Zeng/temporal-robustness-benchmark .
Keywords:
Temporal Action Detection
Temporal Robustness
Video Understanding

Journal

International Journal of Computer Vision cover
International Journal of Computer Vision
IF:
9.3
Papers:
3.9K
Citations:
2.8W

Organization

C
College of Physics and Optoelectronic Engineering
Scholars:
174
Papers: 63
Citations: 1
S
shanghai jiao tong university
Scholars:
15.1W
Papers: 11.5W
Citations: 159
A
artificial intelligence research institute
Scholars:
36
Papers: 20
Citations: 0
C
College of Mechatronics and Control Engineering
Scholars:
53
Papers: 20
Citations: 0
S
state key lab for novel software technology
Scholars:
2
Papers: 1
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers