Return
Enhancing Weakly Supervised Multimodal Video Anomaly Detection Through Text Guidance
S
J
J
龚
DOI:10.1109/tmm.2026.3668927.png)
Abstract
En 中文
In recent years, weakly supervised multimodal video anomaly detection, which leverages RGB, optical flow, and audio modalities, has garnered significant attention from researchers, emerging as a vital subfield within video anomaly detection. However, previous studies have inadequately explored the role of text modality in this domain. With the proliferation of large-scale text-annotated video datasets and the advent of video captioning models, obtaining text descriptions from videos has become increasingly feasible. Text modality, carrying explicit semantic information, can more accurately characterize events within videos and identify anomalies, thereby enhancing the model’s detection capabilities and reducing false alarms. However, text feature extraction challenges anomaly detection. Pre-trained large language models often struggle to effectively capture the nuances associated with anomalies, as their training is based on generalized datasets. Directly fine-tuning the text feature extractor is also challenging, as anomaly-related text descriptions are sparse. Furthermore, due to the varying amounts of information carried by different modalities, issues such as modality redundancy and modality imbalance arise during feature fusion. To address the challenges of text feature extraction and the issues of modality redundancy and imbalance, we propose a novel text-guided weakly supervised multimodal video anomaly detection framework. Specifically, we introduce an in-context learning based multi-stage text augmentation mechanism to generate high-quality anomaly text samples. These high-quality samples are then used to fine-tune the text feature extractor, aiming to obtain a more effective text feature extractor for anomaly detection. Additionally, we present a multi-scale bottleneck Transformer fusion module to enhance multimodal integration, utilizing a set of reduced bottleneck tokens to progressively transmit compressed information across modalities, aiming to address the issues of modality redundancy and imbalance. Experimental results on large-scale datasets UCF-Crime and XD-Violence demonstrate that our proposed approach achieves state-of-the-art performance.
Keywords:
Multimodal video anomaly detection
in-context learning
text augmentation
multi-scale bottleneck transformer
Journal
IF:
9.7
Papers:
4.4K
Citations:
2.4W
