arrow
Return

Frequency-selective countnet: Enhancing text-guided object counting with frequency features

delete2026-01-21
delete0
PRE
AI
C
Cheng Qian
J
Jiwu Cao
Y
Ying Mao
R
Ruotian Zhang
F
Fei Long
J
Jun Sang *
DOI:10.1016/j.patrec.2025.12.014delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Text-guided object counting aims to estimate the number of objects described by natural language within complex visual scenes. However, existing approaches often struggle to align textual intent with diverse visual patterns, especially when target objects vary in scale, appearance, or context. To address these limitations, we propose Frequency-Selective CountNet (FSCNet), a novel framework that integrates spatial and frequency-domain features for precise text-guided counting. FSCNet introduces a Triple-Stream Attention Fusion Module (TSAFM) that combines textual, global, and local visual features. Additionally, an Adaptive Frequency Selector (AFS) dynamically emphasizes frequency components by separately modulating the magnitude and phase spectra, preserving geometric consistency during decoding. Extensive experiments on the FSC-147 and CARPK datasets demonstrate that FSCNet achieves state-of-the-art performance, outperforming previous best methods by 18.34% in MAE and 27.41% in RMSE on FSC-147 (Avg.) and by 5.17%/7.58% on CARPK.
Keywords:
Object counting
Vision-Language model
Frequency domain
Multimodal fusion

Journal

Pattern Recognition Letters cover
Pattern Recognition Letters
IF:
3.3
Papers:
7.8K
Citations:
1.6W

Organization

S
southwest university of science & technology - china
Scholars:
8.5K
Papers: 6.3K
Citations: 6
C
chongqing university
Scholars:
1.2W
Papers: 4.4K
Citations: 1
U
university of glasgow
Scholars:
3.5W
Papers: 3.1W
Citations: 37
A
auckland university of technology
Scholars:
667
Papers: 367
Citations: 0
researcher View more organizations