arrow
Return

Text2Layout: Layout Generation From Text Representation Using Transformer

delete2024-01-01
delete0
delete
OA
AI
H
Haruka Takahashi *
S
Shigeru Kuriyama
DOI:10.1109/ACCESS.2024.3452957delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Recent advanced Text-to-Image methods still require much work to specify all labels and detailed layouts of objects to obtain an accurate planned image. Layout-based synthesis is an alternative method enabling users to control detailed composition directly to avoid the trial-and-error often required in prompt-based editing. This paper proposes a layout generation from text instead of generating images directly. Our approach uses Transformer-based deep neural networks to synthesize scene representations of multiple objects. By focusing on layout information, we can make an explainable layout of what objects the image includes. Our end-to-end approach uses parallel decoding, differs from conventional layout synthesis, and introduces sequential object predictions and post-processing of duplicate bounding boxes. We experimentally compare our method's quality and computational cost against existing ones, demonstrating its effectiveness and efficiency in generating layouts from textual representations. Combined with Layout-to-Image, this approach has significant practical implications, allowing the practical authoring tools that make image generation explainable and computable using relatively lightweight networks.
Keywords:
Layout
Transformers
Decoding
Text to image
Noise measurement
Feature extraction
Text-to-image
layout generation
creation support
transformer

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

T
Toyohashi University of Technology
Scholars:
2.3K
Papers: 1.8K
Citations: 1.5K