Return
Dynamic concept memory and contrastive feature generation for robust generalized zero-shot learning
X
D
M
J
DOI:10.1007/s00371-026-04692-8.png)
Abstract
En 中文
Zero-shot learning (ZSL) aims to recognize unseen classes by transferring semantic knowledge from seen to unseen categories. Recent generative ZSL methods reduce the reliance on human-annotated attributes by using CLIP-based class-name embeddings as semantic conditions. However, such class-name-only semantics mainly encode high-level category information and often lack fine-grained visual cues, leading to insufficient guidance for discriminative feature generation. To address this problem, we propose a dynamic concept memory-guided contrastive generative network (DCM-GAN) for class-name-only generative ZSL. DCM-GAN mines transferable latent visual concepts from seen-class local CLIP tokens and maintains them in a dynamic concept memory to enhance weak class-name semantics. The enhanced semantics are then used to condition a contrastive generative network, where instance-level and category-level contrastive objectives improve the structure of synthesized features. Experiments on AWA2, SUN, CUB, and FLO show that DCM-GAN achieves competitive performance under both CZSL and GZSL settings. Under a controlled CLIP ViT-B/16 setting, DCM-GAN improves the GZSL harmonic mean over reproduced CLIP+TF-VAEGAN by 8.01 and 16.64 percentage points on CUB and SUN, respectively. The source code and reproduction resources are available at: https://github.com/99ww5/DCM-GAN .
Keywords:
Generalized zero-shot learning
Vision–language models
Dynamic concept memory
Contrastive feature generation
Open-world visual recognition
Journal
IF:
2.9
Papers:
4.5K
Citations:
6.5K
