arrow
Return

Multi-Modal Primitive Retrieval for Compositional Zero-Shot Learning

delete2026-02-21
delete0
PRE
AI
C
Chenchen Jing
H
Haozhe Zhang
J
J. G. Lu
Y
Yang Liu
H
Hao Chen
X
Xiaoqin Zhang
C
Chunhua Shen *
DOI:10.1007/s11263-026-02763-wdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Compositional generalization, understanding unseen combinations composed of seen primitives, is one of the fundamental properties of human intelligence. Aiming to evaluate such ability of vision models, compositional zero-shot learning (CZSL) requires recognizing unseen attribute-object compositions by learning from seen compositions. It’s essential for CZSL to compose the learned knowledge of seen primitives, i.e., attributes or objects, into novel compositions. In this work, we propose a retrieval-augmented method to explicitly retrieve knowledge of seen primitives from both vision domain and language domain, for compositional zero-shot learning. Our method augments standard multi-path classification methods with retrieval modules. Specifically, we first construct several databases storing abundant and diverse primitive knowledge, including the attribute and object representations of training images, and textual representations for descriptions of attributes and objects, respectively. For an input training/testing image, we use visual and textual retrieval modules to retrieve representations of relevant training images and text descriptions with the same attribute and object, respectively. The primitive representations and image representation of the input image are augmented by using the retrieved representations, for composition recognition. By referencing semantically similar images and texts, the proposed method is capable of recalling knowledge of seen primitives for compositional generalization. Experiments on three widely used datasets show the effectiveness of the proposed method.
Keywords:
Multi-Modality
Retrieval Augmentation
Compositional Zero-Shot Learning
Compositional Generalization

Journal

International Journal of Computer Vision cover
International Journal of Computer Vision
IF:
9.3
Papers:
3.9K
Citations:
2.8W

Organization

C
college of computer science and technology
Scholars:
302
Papers: 107
Citations: 0
C
cad&cg
Scholars:
3
Papers: 1
Citations: 0
Cited Papers

Cited Papers

Re-ViLM: Retrieval-Augmented Visual Language Model for Zero and Few-Shot Image Captioning
err2023-01-01
err0
PREAI
errYang,Zhuolin; Ping,Wei; Liu,Zihan; Korthikanti,Vijay; Nie,Weili; Huang,De-An; Fan,Linxi; Yu,Zhiding; Lan,Shiyi; Li,Bo; Shoeybi,Mohammad; Liu,Ming-Yu; Zhu,Yuke; Catanzaro,Bryan; Xiao,Chaowei; Anandkumar,Anima
errShare
errSave
Visual question answering: A survey of methods and datasets
err2017-10-01
err0
errOAAI
errQi Wu; Damien Teney; Peng Wang; Chunhua Shen; Anthony Dick; Anton van den Hengel
errShare
errSave
Task-Driven Modular Networks for Zero-Shot Compositional Learning
err2019-10-01
err0
errOAAI
errSenthil Purushwalkam; Maximillian Nickel; Abhinav Gupta; Marc'aurelio Ranzato
errShare
errSave
Reveal: Retrieval-Augmented Visual-Language Pre-Training with Multi-Source Multimodal Knowledge Memory
err2023-06-01
err0
errOAAI
errZiniu Hu; Ahmet Iscen; Chen Sun; Zirui Wang; Kai-Wei Chang; Yizhou Sun; Cordelia Schmid; David A. Ross; Alireza Fathi
errShare
errSave
Estimation of Near-Instance-Level Attribute Bottleneck for Zero-Shot Learning
err2024-02-27
err2
PREAI
errJiang, Chenyi; Shen, Yuming; Chen, Dubing; Zhang, Haofeng; Shao, Ling; Torr, Philip H. S.
errShare
errSave
The Psychology of Associative Learning
err
IF0
err2010-01-26
err0
PREAI
errDavid R. Shanks
errShare
errSave
researcher View more