arrow
Return

Unlocking compositional potential: Leveraging specific feedback for text-to-image generation in diffusion models

delete2025-08-21
delete0
PRE
AI
X
X. Niu *
J
Jinping Tang *
G
Ge Zhu *
DOI:10.1016/j.eswa.2025.129443delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• We construct a novel text-to-image dataset to facilitate research on compositional generation, comprising 11,337 annotated text-image pairs with two key characteristics: 1) Compositional Diversity: covers varying object categories and quantities; 2) Scenario Completeness: includes common, unconventional, and unrealistic scenes, thereby addressing the limitations of existing datasets. • We propose a fine-tuning method with specific feedback for addressing the problem of text-to-image compositional generation, which can be fine-tuned by a differentiable reward function computed by our proposed matching score. • The proposed matching score can also serve as a metric for evaluating text-image alignment in other generation models. • The quantitative and qualitative comparisons demonstrate that our method achieves superior performance over other text-to-image models in both alignment and image quality.
Keywords:
compositional generation
text-to-image synthesis
fine-tuning method
matching score
dataset construction

Journal

Expert Systems with Applications cover
Expert Systems with Applications
IF:
7.5
Papers:
2.9W
Citations:
10.2W

Organization

No organization information available