arrow
Return

Semantically guided dynamic visual prototype refinement for compositional zero-shot learning

delete2026-01-22
delete0
PRE
AI
Z
Zhong Peng
Y
Yishi Xu
G
Gerong Wang
W
Wenchao Chen
J
Jing Zhang *
B
Bo Chen *
H
Hongwei Liu
DOI:10.1016/j.neucom.2026.132775delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Compositional Zero-Shot Learning (CZSL) seeks to recognize unseen state–object pairs by recombining primitives learned from seen compositions. Despite recent progress with vision–language models (VLMs), two limitations remain: (i) text-driven semantic prototypes are weakly discriminative in the visual feature space; and (ii) unseen pairs are optimized passively, thereby inducing seen bias. To address these limitations, we present Duplex, a framework that couples dual-prototype learning with dynamic local-graph refinement of visual prototypes. For each composition, Duplex maintains a semantic prototype via prompt learning and a visual prototype for unseen pairs constructed by recombining disentangled state and object primitives from seen images. The visual prototypes are updated dynamically through lightweight aggregation on mini-batch local graphs, which incorporate unseen compositions during training without labels. This design introduces fine-grained visual evidence while preserving semantic structure. It enriches class prototypes, better disambiguates semantically similar yet visually distinct pairs, and mitigates seen bias. Experiments on MIT-States, UT-Zappos, and CGQA in closed-world and open-world settings achieve competitive performance and consistent compositional generalization.

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

A
Academy of Military Science
Scholars:
146
Papers: 55
Citations: 0
X
Xidian University
Scholars:
2.4W
Papers: 1.9W
Citations: 9.7K