arrow
Return

Extended Graph Learning for Weakly Supervised Video Anomaly Detection

delete2025-10-27
delete0
PRE
AI
J
Jixiang Deng
刘英 cover
刘英 (Ying Liu)
李春光 (Chunguang Li)
DOI:10.1109/TCSVT.2025.3625570delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Video anomaly detection (VAD) is important in many fields because of its theoretical and practical values. One of the challenges in VAD is the difficulty in obtaining segment-level labels due to the high annotation cost. In recent years, researchers have adopted video-level labels as a form of weak supervision, leading to the development of weakly supervised video anomaly detection (WS-VAD). Among different WS-VAD approaches, graph convolutional networks (GCNs) have attracted much attention, since they have the ability to model relationship information in video data. Typically, the relationship, represented by the graph edges, is the class label similarity, and this similarity is built based on the feature similarity and temporal consistency among video segments. Undoubtedly, the more information about class label similarity is provided, the higher the performance of GCN tends to be. In real-world scenarios of VAD, anomalies exhibit several unique properties such as diversity and rarity. These properties may lead to the following situation. Given two video segments, although their feature similarity is low and their time separation is large, both of them are anomalies, that is, they have the same class label. Likewise, normal samples also encounter such situation. However, the existing graph structures in GCN methods do not adequately account for this situation. To address this issue, this paper proposes an extended graph learning (EGL) method that incorporates additional class label similarity among video segments. The proposed EGL includes two extended graph convolutional networks (EGCNs): a spatial EGCN and a temporal EGCN. To capture more accurate information about class label similarity, EGL incorporates a feedback module to update the graph structures of EGCNs. EGL can effectively extract more information about class label similarity, thereby ensuring good performance when training data is scarce. Experimental results highlight the advantages of the proposed EGL method, particularly with limited training samples. In particular, when only 30% of the training data is used, EGL achieves the best performance of 95.55% AUC on ShanghaiTech, 81.29% AUC on UCF-Crime, and 75.31% AP on XD-Violence, outperforming the existing VAD methods by up to 6.86%, 3.23%, and 4.40%, respectively.
Keywords:
Video anomaly detection
graph learning
graph convolutional network
weakly supervised learning

Journal

IEEE Transactions on Circuits and Systems for Video Technology cover
IEEE Transactions on Circuits and Systems for Video Technology
IF:
11.1
Papers:
624
Citations:
3.1W

Organization

Z
zhejiang university
Scholars:
17.6W
Papers: 12.1W
Citations: 152