1
Return

Lightweight Vision Transformer with CBAM and LSTM for Efficient Violence Detection in Image Sequences

delete2026-03-31
delete1
PRE
AI
G
Garg, Aishvarya
N
Nigam, Swati
S
Singh, Rajiv *
DOI:10.1007/s11760-026-05207-7delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In recent years, the need for efficient violence detection methods has become increasingly critical for various security applications. Existing approaches leveraging deep learning (DL), particularly Vision Transformers (ViTs), often suffer from high complexity and overhead, limiting their real-time applicability. This paper proposes a lightweight spatiotemporal transformer-based violence detection method that integrates ViT with a Convolutional Block Attention Module (CBAM) and Long Short-Term Memory (LSTM) for spatial and temporal feature extraction, respectively. By replacing the multi-head self-attention (MSA) mechanism in ViT with CBAM, we significantly reduce the number of parameters and computational costs while enhancing the model's ability to focus on localized patterns essential for violence detection. Our approach is validated on four benchmark datasets, achieving accuracies of 99.81%, 99%, 99.50%, and 99.10% on AIRTLab, Industrial Surveillance, RWF2000, and Hockey Fights, respectively. The proposed method demonstrates a good balance between model complexity and computational efficiency, making it suitable for real-time violence detection applications.
Keywords:
CBAM
Deep learning
LSTM
Violence detection
ViT

Journal

Signal Image and Video Processing cover
Signal Image and Video Processing
IF:
2.1
Papers:
778
Citations:
4.6K

Organization

B
banasthali vidyapith
Scholars:
311
Papers: 148
Citations: 21
Cited Papers

Cited Papers

Citing Papers

Citing Papers