1
Return

Enhancing Weakly Supervised Multimodal Video Anomaly Detection Through Text Guidance

delete2026-03-03
delete0
PRE
AI
S
Shengyang Sun
J
Jiashen Hua
J
Junyi Feng
龚小谨 (Xiaojin Gong)
DOI:10.1109/tmm.2026.3668927delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In recent years, weakly supervised multimodal video anomaly detection, which leverages RGB, optical flow, and audio modalities, has garnered significant attention from researchers, emerging as a vital subfield within video anomaly detection. However, previous studies have inadequately explored the role of text modality in this domain. With the proliferation of large-scale text-annotated video datasets and the advent of video captioning models, obtaining text descriptions from videos has become increasingly feasible. Text modality, carrying explicit semantic information, can more accurately characterize events within videos and identify anomalies, thereby enhancing the model’s detection capabilities and reducing false alarms. However, text feature extraction challenges anomaly detection. Pre-trained large language models often struggle to effectively capture the nuances associated with anomalies, as their training is based on generalized datasets. Directly fine-tuning the text feature extractor is also challenging, as anomaly-related text descriptions are sparse. Furthermore, due to the varying amounts of information carried by different modalities, issues such as modality redundancy and modality imbalance arise during feature fusion. To address the challenges of text feature extraction and the issues of modality redundancy and imbalance, we propose a novel text-guided weakly supervised multimodal video anomaly detection framework. Specifically, we introduce an in-context learning based multi-stage text augmentation mechanism to generate high-quality anomaly text samples. These high-quality samples are then used to fine-tune the text feature extractor, aiming to obtain a more effective text feature extractor for anomaly detection. Additionally, we present a multi-scale bottleneck Transformer fusion module to enhance multimodal integration, utilizing a set of reduced bottleneck tokens to progressively transmit compressed information across modalities, aiming to address the issues of modality redundancy and imbalance. Experimental results on large-scale datasets UCF-Crime and XD-Violence demonstrate that our proposed approach achieves state-of-the-art performance.
Keywords:
Multimodal video anomaly detection
in-context learning
text augmentation
multi-scale bottleneck transformer

Journal

IEEE Transactions on Multimedia cover
IEEE Transactions on Multimedia
IF:
9.7
Papers:
4.4K
Citations:
2.4W

Organization

A
alibaba cloud
Scholars:
39
Papers: 10
Citations: 1
H
Hangzhou Dianzi University
Scholars:
1.2W
Papers: 9.4K
Citations: 7.5K
Z
zhejiang university
Scholars:
17.0W
Papers: 11.9W
Citations: 152
Cited Papers

Cited Papers

Citing Papers

Citing Papers