Return
Discriminative context fusion network for crowd counting in complex park environment☆
Z
Z
W
W
M
S
DOI:10.1016/j.displa.2026.103434.png)
Abstract
En 中文
To address the challenges of insufficient cross-modal information fusion and difficult multi-scale context coordination in crowd counting tasks under complex industrial park environments, a novel Discriminative Context Fusion Network (DCFNet) was proposed. The network first realizes bi-directional and adaptive excitation and enhancement for RGB and thermal features through a designed Bi-directional Discriminative Excitation Module (BDEM), effectively solving the problem of dynamic variations in modal information value under different lighting conditions. Secondly, to coordinate multi-scale information more precisely, a Frequency-Gated Context Fusion Module (FGCFM) was constructed. It decouples features into low-frequency and high-frequency components in the frequency domain and utilizes adjacent-level context for targeted gating enhancement to adapt to significant crowd scale variations within the park. Experimental results on the public RGBT-CC dataset show that the performance of the proposed method is superior to that of various existing state-of-the-art methods.
Keywords:
RGB-T crowd counting
Industrial park
Cross-modal interaction
Vision-language pre-training
Journal
IF:
3.4
Papers:
2.1K
Citations:
3.2K
