arrow
Return

Diffusion-Based Text-Guided Image Generation With Fine-Grained Spatial Object-Attribute Relationships

delete2025-12-01
delete0
PRE
AI
F
Fuxiang Wu
L
Liu Liu
F
Fusheng Hao
Z
Ziliang Ren
D
Dacheng Tao
X
Xinyu Wu
J
Jun Cheng
DOI:10.1109/TCSVT.2025.3639218delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Expressing and controlling fine-grained spatial attributes of objects in large-scale models presents significant challenges, as these spatial attributes are often difficult to describe textually and exhaustive enumeration is impractical. This hinders effective alignment with user preferences regarding spatial attribute-object relationships in fine-grained synthesis tasks. To tackle this problem, we propose AttrObjDiff, a novel framework built on the pre-trained Stable Diffusion model to integrate spatial attribute maps. Firstly, AttrObjDiff constrains the denoising step using trainable cross-attention fusion modules, attribute-enhancing cross-attention and LoRAs. The fusion modules take layout features extracted by a frozen ControlNet and corresponding fine-grained attribute maps as inputs to generate joint constraint features of spatial attribute-object relationships. We leverage attribute-enhancing cross-attention within the U-Net to further refine these spatial attributes. Finally, LoRAs are employed to align with these joint constraint features of fine-grained relationships. Secondly, AttrObjDiff enhances the reverse process with lightweight noise reranking models to improve spatial object-attribute alignment. The reranking models select semantic noises related to fine-grained relationships, improving synthesis quality without significantly increasing computational costs. Experimental results demonstrate that our method can generate high-quality images guided by fine-grained spatial object-attribute relationships, improving synthesis controllability and semantic consistency.
Keywords:
Attribute map
lightweight noise reranking
cross-attention fusion

Journal

IEEE Transactions on Circuits and Systems for Video Technology cover
IEEE Transactions on Circuits and Systems for Video Technology
IF:
11.1
Papers:
612
Citations:
3.1W

Organization

B
beihang university
Scholars:
5.2K
Papers: 2.0K
Citations: 21
N
nanyang technological university
Scholars:
2.5K
Papers: 1.6K
Citations: 1
D
dongguan university of technology
Scholars:
848
Papers: 330
Citations: 0
C
Chinese Academy of Sciences
Scholars:
3.9W
Papers: 1.5W
Citations: 58.4W
C
chinese academy of sciences
Scholars:
55.9W
Papers: 44.7W
Citations: 704
researcher View more organizations