arrow
Return

A Customizable Face Generation Method Based on Stable Diffusion Model

delete2024-01-01
delete0
delete
OA
AI
W
Wenlong Xiang
S
Shuzhen Xu *
C
Cuicui Lv
S
Shuo Wang
DOI:10.1109/ACCESS.2024.3520719delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Facial generation technology uses computer algorithms and artificial intelligence techniques to generate realistic facial images. This technology typically employs deep learning models such as Generative Adversarial Networks (GANs) and diffusion models, learning from a large dataset of real facial images to generate new virtual facial images. Text-to-image model combines textual and visual information by generating image content based on corresponding input text descriptions. The development of text-to-image models provides new insights into cross-modal learning, promotes interaction and fusion between textual and visual information, thereby supporting the advancement of multimodal intelligent systems. This paper aims to apply text-to-image models to design a customizable facial generation model. The model is improved upon the Stable Diffusion model by incorporating LoRA (Low-Rank Adaptation) principles for style constraints. In addition, we modify the structure of Variational Autoencoder to enhance generation efficiency. Upon obtaining initial generation results, local refinements can be performed without altering the main structure, allowing for localized adjustments. Through pre-generation and post-processing, our model can accurately generate semantically guided face images.
Keywords:
Diffusion models
Noise
Computational modeling
Adaptation models
Noise reduction
Training
Text to image
Semantics
Image synthesis
Computer architecture
Diffusion model
facial image generation
text-to-image
U-Net

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

Y
Yantai University
Scholars:
8.4K
Papers: 5.7K
Citations: 9.9K