arrow
Return

Graph visual representation for controllable scene layout generation

delete2026-05-08
delete0
PRE
AI
李津 cover
李津 (Jin Li)
M
Minghan Ma
L
Longjiang Guo *
X
Xue Dong
M
Meirui Ren *
DOI:10.1016/j.jvcir.2026.104814delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Natural scene layouts and object bounding boxes are essential for controllable text-to-image generation. Diffusion-based methods often exhibit object misplacements and scaling inaccuracies due to insufficient explicit modeling of hierarchical spatial relationships. Existing layout generation approaches primarily encode descriptive triplets into sentences, overlooking explicit graphic signals. To address this, we propose GCN-LT, a novel method employing Graph Convolutional Networks (GCN) to explicitly capture object relationship features, which are then fused with implicit semantic features. A Transformer encoder–decoder processes these fused features to predict bounding boxes compliant with semantic and spatial constraints. Evaluations on COCO and VG datasets demonstrate that GCN-LT outperforms state-of-the-art baselines in generating natural scene layouts, resulting in more natural and coordinated images. GCN-LT provides an effective solution for intelligent layout generation.
Keywords:
Graph Convolutional Networks
Scene Layout Generation
Object Bounding Boxes
Controllable Text-to-Image Generation
Visual Representation

Journal

Journal of Visual Communication and Image Representation cover
Journal of Visual Communication and Image Representation
IF:
3.1
Papers:
414
Citations:
5.6K

Organization

S
shaanxi normal university
Scholars:
1.9K
Papers: 600
Citations: 0