arrow
Return

Sketch-guided scene image generation with diffusion model

delete2025-06-01
delete0
PRE
AI
T
Tianyu Zhang
X
Xiaoxuan Xie
X
Xusheng Du
H
Haoran Xie *
DOI:10.1016/j.cag.2025.104226delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Text-to-image models showcase the impressive ability to generate high-quality and diverse images. However, the transition from freehand sketches to complex scene images with multiple objects remains challenging in computer graphics. In this study, we propose a novel sketch-guided scene image generation framework, decomposing the task of scene image generation from sketch inputs into object-level cross-domain generation and scene-level image construction steps. We first employ a pre-trained diffusion model to convert each single object drawing into a separate image, which can infer additional image details while maintaining the sparse sketch structure. To preserve the conceptual fidelity of the foreground during scene generation, we invert the visual features of object images into identity embeddings for scene generation. For scene-level image construction, we generate the latent representation of the scene image using the separated background prompts. Then, we blend the generated foreground objects with the background image guided by the layout of sketch inputs. We infer the scene image on the blended latent representation using a global prompt with the trained identity tokens to ensure the foreground objects' details remain unchanged while naturally composing the scene image. Through qualitative and quantitative experiments, we demonstrated that the proposed method's ability surpasses the state-of-the-art approaches for scene image generation from hand-drawn sketches.
Keywords:
Sketch-guided
Scene image
Generative model
Diffusion model

Journal

C
Computers and Graphics
IF:
2.8
Papers:
82
Citations:
4.3K

Organization

J
Japan Advanced Institute of Science and Technology
Scholars:
283
Papers: 169
Citations: 1.8K