Return
Exploring What Why and How: A Multifaceted Benchmark for Causation Understanding of Video Anomaly
H
G
J
W
W
H
Z
J
M
P
J
X
J
DOI:10.1007/s11263-026-02983-0.png)
Abstract
En 中文
Recent progress in video anomaly understanding (VAU) enables significant applications in traffic monitoring and industrial automation, yet existing benchmarks mainly focus on anomaly detection and localization. We push VAU toward practical comprehension by explicitly evaluating whether models can describe what anomaly occurred, explain why it happened, and infer what effect it caused. To this end, we introduce ECVA, a benchmark for Exploring the Causation of Video Anomalies, where each video is paired with detailed human annotations covering (1) anomaly type, temporal boundaries, and event descriptions, (2) natural-language explanations of causes, and (3) free-form descriptions of effects. In addition, ECVA provides annotation-derived importance curves to characterize the relative contribution of key evidence segments, supporting fine-grained evaluation and reliability analysis. Building on ECVA, we propose AnomShield, a video large language model for reasoning-intensive anomaly understanding. AnomShield adopts a Chain-of-Thought reasoning paradigm to explicitly determine and extract anomaly-relevant temporal segments, and subsequently employs spatiotemporal-decoupled positional encoding together with bi-directional state-space modeling to capture fine-grained spatiotemporal dependencies. To evaluate models under ECVA’s setting, we introduce AnomEval, a human-aligned metric for more reliable assessment of video-LLMs. Extensive experiments validate the effectiveness of our benchmark, model, and metric. Code and dataset are available at https://github.com/Dulpy/ECVA .
Keywords:
Video anomaly understanding
Causal reasoning
Video Large language model
Evaluation metric
Journal
IF:
9.3
Papers:
3.9K
Citations:
2.8W
