Return
Transformer-enhanced lightweight detector: YOLOv11-TSO for small objects in traffic surveillance
Y
Y
DOI:10.1016/j.aej.2026.05.033.png)
Abstract
En 中文
Camera-based small object detection is essential for intelligent traffic surveillance, yet existing YOLO variants, Transformer-enhanced detectors, and lightweight traffic detection models still struggle to jointly satisfy detection accuracy, real-time inference, and edge deployment requirements under complex working conditions. To address this problem, this paper proposes YOLOv11-TSO, a Transformer-enhanced lightweight detector specifically designed for traffic small object detection in adverse environments. The central objective is to improve the detection of distant pedestrians, compact vehicles, bicycles, and partially occluded targets while maintaining practical inference efficiency on resource-constrained devices. YOLOv11-TSO introduces three coordinated innovations: a C2PSA-SA backbone that embeds Transformer self-attention for lightweight global contextual modeling, a C3k2_GAM neck that integrates group attention and refined multi-scale feature fusion, and label smoothing loss to mitigate overfitting and class imbalance. Extensive experiments on VisDrone2019 and WiderPerson, together with robustness tests under fog, rain, occlusion, and low-light conditions, demonstrate the effectiveness of the proposed model. On VisDrone2019, YOLOv11-TSO achieves 93.7% precision, 96.1% recall, and 92.4% mAP@0.5 with only 9.2M parameters and 34.0 GFLOPs. It reaches 35 FPS on NVIDIA H100, 18 FPS on Jetson TX2, and 8 FPS on Jetson Nano, indicating a practical accuracy–efficiency–deployment trade-off for intelligent traffic surveillance. These results confirm that YOLOv11-TSO provides a robust and lightweight solution for small object detection in complex traffic environments.
Keywords:
YOLOv11-TSO
Small object detection
Transformer self-attention mechanism
Traffic monitoring
Complex environments
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
6.8
Papers:
6.3K
Citations:
2.6W
