arrow
Return

Position-Aware Text-to-Image Generation with Efficient Controllability

delete2026-01-01
delete0
PRE
AI
G
Gu, Junchao *
王翔宇 cover
王翔宇 (Xiangyu Wang)
Y
Yuchen Du
H
Hao Chen
DOI:10.1007/978-981-95-3398-5_12delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In recent years, Stable Diffusion has significantly advanced the quality of text-to-image generation. However, accurately interpreting and representing Spatial Layouts specified by text prompts remains challenging. Existing approaches typically rely either on extra grounding information or utilize LLMs (large language models) combined with layout-controllable models which suffer from high computational costs. To address these limitations, we propose a novel method SamLayGe, a lightweight and efficient layout generation model designed for seamless integration into existing text-to-image pipelines. SamLayGe autonomously generates comprehensive layouts without requiring explicit user inputs, thus surpassing current LLM-based and layout-controllable approaches in terms of versatility and efficiency. Furthermore, we propose LayGeBench, a benchmark dataset addressing ambiguities in spatial descriptions of prior datasets. Extensive evaluations demonstrate that SamLayGe consistently produces images that accurately adhere to textual layout descriptions, achieving superior performance in terms of both accuracy and computational efficiency. Code is available at https://github.com/shenlanzhuanshu/caption-to-positional-layout. Position-Aware Text-to-Image Generation with Efficient Controllability.
Keywords:
Text-to-image
Text-to-layout
Efficient Controllability
Positional Relationship

Journal

I
IMAGE AND GRAPHICS, ICIG 2025, PT I
IF:
0
Papers:
38
Citations:
0

Organization

X
xidian university
Scholars:
5.8K
Papers: 2.0K
Citations: 0