arrow
Return

Attribute-Centric Compositional Text-to-Image Generation

delete2025-03-13
delete0
delete
OA
AI
Y
Yuren Cong
M
Martin Renqiang Min
L
Li Erran Li
B
Bodo Rosenhahn
M
Michael Ying Yang *
DOI:10.1007/s11263-025-02371-0delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Despite the recent impressive breakthroughs in text-to-image generation, generative models have difficulty in capturing the data distribution of underrepresented attribute compositions while over-memorizing overrepresented attribute compositions, which raises public concerns about their robustness and fairness. To tackle this challenge, we propose ACTIG, an attribute-centric compositional text-to-image generation framework. We present an attribute-centric feature augmentation and a novel image-free training scheme, which greatly improves model's ability to generate images with underrepresented attributes. We further propose an attribute-centric contrastive loss to avoid overfitting to overrepresented attribute compositions. We validate our framework on the CelebA-HQ and CUB datasets. Extensive experiments show that the compositional generalization of ACTIG is outstanding, and our framework outperforms previous works in terms of image quality and text-image consistency. The source code and trained models are publicly available at https://github.com/yrcong/ACTIG.
Keywords:
Text-to-image
Compositional generation
Attribute-centric

Journal

International Journal of Computer Vision cover
International Journal of Computer Vision
IF:
9.3
Papers:
3.9K
Citations:
2.8W

Organization

N
nec labs amer
Scholars:
1
Papers: 1
Citations: 0
U
Univ Bath
Scholars:
481
Papers: 421
Citations: 107
A
amazon.com
Scholars:
698
Papers: 505
Citations: 8
researcher View more organizations