arrow
Return

OMR-diffusion: Optimizing multi-round enhanced training in diffusion models for improved intent understanding

delete2025-10-09
delete0
PRE
AI
K
Kun Li
J
Jianhui Wang
Y
Yangfan He
M
Miao Zhang *
王学谦 (Xueqian Wang)
DOI:10.1016/j.neucom.2025.131452delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• Optimizing User Preference in Image Generation: This work introduces a Visual Co-Adaptation (VCA) framework that integrates human feedback and reinforcement learning (RL) to align generated images with personalized user preferences while maintaining consistency across multiple turns. • Improving Generation Through Multi-Round Feedback: Dialogue-based refinement leverages user feedback at each step to enhance image diversity, structural consistency, and semantic alignment, significantly advancing text-to-image generation. • Enhancing Interaction with Feedback Loops: A dynamic reward system balances diversity, consistency, and mutual information to improve interaction quality and lower the barriers to utilizing AI technology. • Exceeding State-of-the-Art Performance: The proposed model surpasses advanced systems like DALL-E 3 and Imagen in intent alignment, image consistency, and dialogue efficiency, achieving top results in LPIPS (0.15) and BLIP (0.59).

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

U
university of electronic science and technology of china
Scholars:
1.3W
Papers: 4.7K
Citations: 4
T
tsinghua university
Scholars:
11.8W
Papers: 10.0W
Citations: 137
X
xiamen university
Scholars:
5.8W
Papers: 3.8W
Citations: 67
researcher View more organizations