Return
Image-Attribute and Frequency-Spatial Dual Collaborative Learning for Pedestrian Attribute Recognition
DOI:10.1109/TIFS.2025.3622074.png)
Abstract
En 中文
The key of pedestrian attribute recognition (PAR) task lies in accurately extracting various attributes from pedestrian images, such as gender, age, clothing, and accessories. Prior CLIP-based PAR methods primarily rely on static textual prompts or template-based sentences and perform direct classification or one-to-one image-text matching. Although these methods have achieved promising results, their reliance on such static prompts may lack flexibility in capturing the dynamic interactions between multiple attribute labels and their fine-grained visual semantics. To address this, we propose a novel Frequency-Spatial and Image-Attribute (FSIA) method that adopts a dual collaborative learning strategy to effectively model image-attribute associations and enhance feature representation. FSIA consists of two key components: 1) An Image-Attribute Collaborative Learning (IACL) framework that integrates visual data with attribute labels to enable a nuanced semantic understanding. The proposed IACL utilizes learnable attribute prompts, each specifically optimized for an individual attribute category, facilitating expressive and discriminative visual-language alignment; 2) A Frequency-Spatial Collaborative Learning (FSCL) module that leverages frequency-domain information (often overlooked in prior PAR) to exploit dual-domain frequency-spatial information in pedestrian images, thereby enhancing feature robustness. Extensive experiments show that FSIA has significant advantages in improving the performance of PAR and the challenging zero-shot PAR task. Specifically, the proposed FSIA achieves 90.08% in mA, 82.09% in Accu, 87.92% in Prec, 89.73% in Recall and 88.45% in F1 on the PETA dataset.
Keywords:
Frequency domain
spatial domain
dual collaborative learning
pedestrian attribute recognition
prompts
Journal
IF:
8
Papers:
5.2K
Citations:
2.3W

