返回
Multi2Human: Controllable human image generation with multimodal controls
DOI:10.1016/j.neucom.2024.127682.png)
摘要
En 中文
Generating high -quality and diverse human images presents a substantial difficulty within the field of computer vision, especially in developing controllable generative models that can utilize input from various modalities. Such models could enable innovative applications like digital human, fashion design, and content creation. In this study, we introduce Multi2Human, a two -stage image synthesis framework for controllable human image generation with multimodal controls . In the first stage, a novel WAvelet-vqVaE (WAVE) architecture is designed to embed human images using a learnable codebook. The WAVE model enhances the conventional Vector Quantized Variational Autoencoder (VQVAE) by integrating wavelets throughout the encoder, thereby enhancing the quality of image reconstruction and synthesis. In the second stage, a new Multimodal Conditioned Diffusion Model (MCDM) is designed to estimate the underlying prior distribution within the discrete latent space using a discrete diffusion process, thus allowing for human image generation conditioned on multimodal controls. Quantitative and qualitative analysis demonstrates that the proposed method has the ability to create high -quality, lifelike full -body human images while satisfying the specified multimodal controls. Our code is available at https://github.com/gxl-groups/Multi2Human.
Keyword:
Controllable image generation
Human image
Diffusion model
Multimodal guidance
期刊
IF:
6.5
论文数:
2.5W
被引数:
6.5W
机构
引用论文
Sensorless Control of Z Source Inverter fed BLDC Motor Drive by FOC - DTC Hybrid Control Strategy Using Fuzzy Logic Controller采用模糊逻辑控制器的foc-dtc混合控制策略的Z源逆变器馈电BLDC电机驱动的无传感器控制
PCCM-GAN: Photographic Text-to-Image Generation with Pyramid Contrastive Consistency Model
NEUROCOMPUTING
IF6.5

