Return
Unlocking compositional potential: Leveraging specific feedback for text-to-image generation in diffusion models
DOI:10.1016/j.eswa.2025.129443.png)
Abstract
En 中文
• We construct a novel text-to-image dataset to facilitate research on compositional generation, comprising 11,337 annotated text-image pairs with two key characteristics: 1) Compositional Diversity: covers varying object categories and quantities; 2) Scenario Completeness: includes common, unconventional, and unrealistic scenes, thereby addressing the limitations of existing datasets. • We propose a fine-tuning method with specific feedback for addressing the problem of text-to-image compositional generation, which can be fine-tuned by a differentiable reward function computed by our proposed matching score. • The proposed matching score can also serve as a metric for evaluating text-image alignment in other generation models. • The quantitative and qualitative comparisons demonstrate that our method achieves superior performance over other text-to-image models in both alignment and image quality.
Keywords:
compositional generation
text-to-image synthesis
fine-tuning method
matching score
dataset construction
Journal
IF:
7.5
Papers:
2.9W
Citations:
10.2W
Organization
No organization information available

