Return
MSLAN: multi-scale spatial location awareness network for crowd counting
DOI:10.1007/s10044-026-01748-2.png)
Abstract
En 中文
Crowd counting, a fundamental task in computer vision, remains challenging due to the simultaneous presence of severe scale variations and complex background noise in real-world scenes. To address these issues, we propose a Multi-scale Spatial Location Awareness Network (MSLAN), which uses MPViT-small as the backbone and introduces Multi-scale Group-wise Attention (MGA) for multi-scale feature representation. Unlike conventional multi-scale designs that mainly enlarge receptive fields, MGA performs attention within channel groups. By integrating spatial pooling, local convolution, and channel projection, it enables different channel subspaces to capture scale-specific crowd patterns, thereby reducing semantic interference across scales. Furthermore, the Spatial Location Awareness Module enhances robustness through random masking of high-level semantic features, adaptive multi-scale feature fusion, and coordinate-enhanced attention mechanisms that strengthen responses to head regions. This design is particularly suited to crowd counting, as it enables the network to focus on reliable head-region cues under severe occlusions and background clutter. Extensive experiments conducted on the ShanghaiTech, UCF-QNRF, JHU-Crowd++, and NWPU-Crowd datasets demonstrate that MSLAN achieves competitive counting accuracy while maintaining a favorable trade-off between performance and computational efficiency. Code is available at https://github.com/qiqi304/MSLAN .
Keywords:
Crowd counting
Spatial location awareness
Scale variations
Background noise
Journal
IF:
2
Papers:
1.9K
Citations:
1.9K

