arrow
Return

Attribute Prompt Alignment Network for Zero-Shot Learning

delete2025-08-19
delete0
PRE
AI
G
Guo-Sen Xie
J
Junyi Li
T
Ting Guo
X
Xiangbo Shu
F
Fang Zhao
张政 cover
张政 (Zheng Zhang)
L
Ling Shao
DOI:10.1109/TNNLS.2025.3598191delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In the vanilla zero-shot learning (ZSL) paradigm, category <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">attributes</i> is the key for knowledge generalizable transfer from seen to unseen classes. By contrast, the current contrastive language-image pretraining (CLIP) model relies on the category <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">names</i> to achieve a more general ZSL-like prediction. When vanilla ZSL meets general CLIP, however, most existing methods on both sides struggle to benefit from each other. In this brief, we resort to attribute prompt tuning (APT) for improving the knowledge transferability from the pretrained CLIP model to the downstream ZSL framework for pursuing desirable feature representations. Our approach, termed as attribute prompt alignment network (APAN), leverages APT for cross-network feature alignment (CFA). In this way, we can investigate the effects of CLIP to vanilla ZSL task in the era of large model by the two branch APAN architecture. Specifically, APT takes as an input the templates of class attribute descriptions to produce attribute prompts, which are further used to both guide the localizations of visual regions across two frozen feature extraction networks, through a visual-semantic interaction attention. This enables APAN to progressively refine and align these cross-network features, thus resulting in generalizable feature representations that can capture fine-grained attribute information. For CFA, we simply introduce prediction alignment loss that constrains the predictions from these two cross-network visual features. Experimental results on three benchmark datasets well demonstrate that APAN outperforms the state-of-the-art methods by absorbing generalizable knowledge from CLIP models.
Keywords:
Attention
attribute prompt
contrastive language-image pretraining (CLIP)
zero-shot learning (ZSL)

Journal

IEEE Transactions on Neural Networks and Learning Systems cover
IEEE Transactions on Neural Networks and Learning Systems
IF:
8.9
Papers:
7.5K
Citations:
7.2W

Organization

N
Nanjing University of Science and Technology
Scholars:
5.6K
Papers: 2.2K
Citations: 25
N
North University of China
Scholars:
1.1W
Papers: 6.9K
Citations: 7.7K
H
harbin institute of technology
Scholars:
8.0W
Papers: 6.6W
Citations: 66
U
University of Chinese Academy of Sciences
Scholars:
6.0K
Papers: 2.4K
Citations: 24.6W
N
nanjing university
Scholars:
7.7W
Papers: 5.6W
Citations: 87
researcher View more organizations