Return
Graph visual representation for controllable scene layout generation
DOI:10.1016/j.jvcir.2026.104814.png)
Abstract
En 中文
Natural scene layouts and object bounding boxes are essential for controllable text-to-image generation. Diffusion-based methods often exhibit object misplacements and scaling inaccuracies due to insufficient explicit modeling of hierarchical spatial relationships. Existing layout generation approaches primarily encode descriptive triplets into sentences, overlooking explicit graphic signals. To address this, we propose GCN-LT, a novel method employing Graph Convolutional Networks (GCN) to explicitly capture object relationship features, which are then fused with implicit semantic features. A Transformer encoder–decoder processes these fused features to predict bounding boxes compliant with semantic and spatial constraints. Evaluations on COCO and VG datasets demonstrate that GCN-LT outperforms state-of-the-art baselines in generating natural scene layouts, resulting in more natural and coordinated images. GCN-LT provides an effective solution for intelligent layout generation.
Keywords:
Graph Convolutional Networks
Scene Layout Generation
Object Bounding Boxes
Controllable Text-to-Image Generation
Visual Representation
Journal
IF:
3.1
Papers:
414
Citations:
5.6K

