Return
Explicit knowledge-structured weakly supervised video anomaly detection
DOI:10.1016/j.dsp.2026.106088.png)
Abstract
En 中文
Video anomaly detection (VAD) is crucial for safety-critical applications such as smart-city surveillance and autonomous driving. While weakly supervised video anomaly detection (WSVAD) reduces annotation cost by using video-level labels, many existing methods rely on a rigid top-k snippet selection strategy that can introduce noisy supervision and overlook explicit prior knowledge required for inferring the time of occurrence, location, and semantic category of an anomaly. Accordingly, we propose Explicit Knowledge-Structured Weakly Supervised Video Anomaly Detection (EKS-WSVAD), which adaptively mines high-confidence anomalous snippets using batch-wise statistics instead of a fixed top-k quota. Moreover, we design a two-stream dynamic-static architecture to incorporate explicit multimodal knowledge for joint reasoning over motion cues, semantic categories, and spatial layouts. Experiments show that EKS-WSVAD achieves 86.76% AUC on UCF-Crime and 85.30% AP on XD-Violence, outperforming most state-of-the-art methods.
Keywords:
Video anomaly detection
Weakly supervised learning
Anomalous snippet mining
Multimodal knowledge
Dynamic-static architecture
Journal
D
IF:
3
Papers:
653
Citations:
0

