Return
Text-guided patch-level exemplar selection for zero-shot counting
DOI:10.1016/j.engappai.2026.113986.png)
Abstract
En 中文
Text-guided zero-shot object counting aims to estimate object counts in images via text prompts. Existing methods face two key challenges. Firstly, current exemplar modeling methods overlook cross-modal fusion, whether using text as exemplars or selecting visuals based on text. Secondly, image–exemplar feature interaction suffers from semantic shift, causing models to focus on incorrect regions. To address the above issues, we propose a text-guided patch-level exemplar selection counting method (TPECount), which is based on Contrastive Language–Image Pretraining (CLIP). It mitigates the absence of visual appearance details while preserving textual semantic information in selected exemplar features and generates high-quality correlation features for density map regression. Specifically, TPECount comprises two well-designed modules: Visual Exemplar Optimization Module (VEOM) and Similarity-Guided Interaction Module (SGIM). VEOM leverages cross-modal feature fusion between images and text to model and select patch-level exemplar features, enabling TPECount to simultaneously leverage relevant exemplar visual appearance features and text features. SGIM is a novel interaction module that consists of our well-designed multi-layer transformer decoder blocks for image–exemplar feature interaction. By leveraging learnable image–exemplar similarity as global supervision, SGIM mitigates the impact of semantic shift and compute correlation features. Extensive experiments on three counting datasets demonstrate that our method achieves superior performance and generalizability. The code is available at https://github.com/aHui3/TPECount .
Journal
IF:
8
Papers:
5.3K
Citations:
3.5W

