arrow
Return

Multi-Object Tracking and Spatio-Temporal Graph-Based Relational Pattern Mining

delete2026-09-11
delete0
delete
OA
AI
T
Taeyoung Seo
K
Kyungyong Chung *
DOI:10.3390/electronics15184086delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This paper addresses the task of binary violence classification in short surveillance video clips, i.e., deciding whether a given clip contains violent interactions (Fight) or not (Non-Fight). Although this task is often discussed in the broader context of video-based abnormal behavior detection driven by smart cities, Closed-Circuit Television (CCTV) networks, and intelligent surveillance systems, existing methods for violence classification largely rely on appearance features from single frames or global motion cues across whole scenes and therefore fail to adequately capture the interactions among multiple objects and the structural changes in their relationships that arise in real surveillance environments. In particular, violent behavior is rarely defined by a specific pose or a single moment; rather, it emerges as a cumulative process in which relational changes—such as inter-person approach, distance variation, collision, and repeated contact—unfold over time. Detecting such behavior accurately therefore requires an approach that can analyze inter-object relationships in a spatiotemporal manner. To this end, this paper proposes a violence detection method that combines multi-object tracking with spatiotemporal graph-based relational pattern mining. The proposed method first detects and tracks person objects using YOLO and DeepSORT, and extracts time-series features—including position, velocity, pose, and inter-object distance variation—to construct a spatiotemporal graph. Relational event sequences are then generated from the edge features of the graph, and class-representative relational patterns are automatically extracted based on discriminative power through PrefixSpan-based frequent sequential pattern mining. In parallel, the spatiotemporal graph is fed into a Spatial Temporal Graph Convolutional Network (ST-GCN) to learn the structural relationships among objects and their temporal evolution. Finally, the pattern-matching score and the ST-GCN classification score are combined to classify each input video as either violent or non-violent. By jointly exploiting interpretable relational pattern information and graph-based structural learning, the proposed approach compensates for the limitations of appearance-centric anomaly detection and demonstrates its applicability to complex real-world surveillance environments. Performance is evaluated in terms of Accuracy, Precision, Recall, and F1-score, with Recall considered a primary metric to reflect the importance of not missing violent events.
Keywords:
violence detection
multi-object tracking
spatiotemporal graph
relational pattern mining
graph neural network
ST-GCN

Journal

Electronics cover
Electronics
IF:
2.6
Papers:
9.8K
Citations:
4.7W

Organization

K
kyonggi university
Scholars:
332
Papers: 231
Citations: 0
Cited Papers

Cited Papers

CNN features with bi-directional LSTM for real-time anomaly detection in surveillance networks
err2020-08-20
err139
PREAI
errUllah, Waseem; Ullah, Amin; Ul Haq, Ijaz; Muhammad, Khan; Sajjad, Muhammad; Baik, Sung Wook
errShare
errSave
A Survey of Video Surveillance Systems in Smart City
err2023-08-23
err0
errOAAI
errYanjinlkham Myagmar-Ochir; Wooseong Kim
errShare
errSave
DeepReS: A Deep Learning-Based Video Summarization Strategy for Resource-Constrained Industrial Surveillance Scenarios
err2020-09-01
err56
PREAI
errMuhammad, Khan; Hussain, Tanveer; Del Ser, Javier; Palade, Vasile; de Albuquerque, Victor Hugo C.
errShare
errSave
researcher View more