arrow
Return

Query-Based Knowledge Sharing for Open-Vocabulary Multi-Label Classification

delete2025-12-01
delete0
delete
OA
AI
X
Xuelin Zhu
J
Jian Liu
D
Dongqi Tang
J
Jiawei Ge
W
Weijia Liu
刘波 cover
刘波 (Bo Liu)
曹玖新 (Jiuxin Cao) *
DOI:10.1145/3762195delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Identifying labels that are unseen during training, known as multi-label zero-shot learning, is a non-trivial task in computer vision. Recent studies have increasingly focused on utilizing vision-language pre-training (VLP) models to recognize unseen labels in an open-vocabulary manner. However, these approaches like knowledge distillation have offered only modest performance gains. The challenge of fully harnessing the potential of VLP models for effective multi-label zero-shot learning remains open. In this work, an advanced query-based knowledge sharing framework is proposed to explore the multi-modal knowledge from VLP models for open-vocabulary multi-label classification. Specifically, we introduce a set of label-agnostic query tokens that are designed to capture essential and informative visual knowledge from input images. These tokens are subsequently shared across all labels, allowing them to select pertinent one as visual clues for accurate recognition. Then, by integrating the pre-trained knowledge of VLP models, these query tokens, trained on seen labels, can be efficiently generalized to the recognition of unseen labels. Additionally, we reformulate ranking learning into a form of classification to enable the magnitude of feature vectors for prediction, which significantly benefits label recognition. Experiment results show that our framework outperforms state-of-the-art methods in multi-label zero-shot learning task by a significant margin, reaching 4.2% and 2.4% in mAP on the NUS-WIDE and Open Images datasets, respectively. Code and models are available at https://github.com/jasonseu/QKS.
Keywords:
Multi-label image classification
Open-vocabulary classification
Vision-language pre-training
CLIP
Prompt learning

Journal

ACM Transactions on Multimedia Computing Communications and Applications cover
ACM Transactions on Multimedia Computing Communications and Applications
IF:
6
Papers:
2.0K
Citations:
5.4K

Organization

S
southeast university - china
Scholars:
5.3W
Papers: 4.9W
Citations: 57
A
Ant Group
Scholars:
23
Papers: 10
Citations: 0