arrow
Return

Multi3D: 3D-aware multimodal image synthesis

delete2024-04-03
delete1
delete
OA
AI
W
Wenyang Zhou
L
Lu Yuan
穆太江 (Tai‐Jiang Mu) *
DOI:10.1007/s41095-024-0422-4delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
3D-aware image synthesis has attained high quality and robust 3D consistency. Existing 3D controllable generative models are designed to synthesize 3D-aware images through a single modality, such as 2D segmentation or sketches, but lack the ability to finely control generated content, such as texture and age. In pursuit of enhancing user-guided controllability, we propose Multi3D, a 3D-aware controllable image synthesis model that supports multi-modal input. Our model can govern the geometry of the generated image using a 2D label map, such as a segmentation or sketch map, while concurrently regulating the appearance of the generated image through a textual description. To demonstrate the effectiveness of our method, we have conducted experiments on multiple datasets, including CelebAMask-HQ, AFHQ-cat, and shapenet-car. Qualitative and quantitative evaluations show that our method outperforms existing state-of-the-art methods.
Keywords:
generate adversarial networks (GANs)
neural radiation field (NeRF)
3D-aware image synthesis
controllable generation

Journal

Computational Visual Media cover
Computational Visual Media
IF:
18.3
Papers:
310
Citations:
2.6K

Organization

T
tsinghua university
Scholars:
11.7W
Papers: 10.0W
Citations: 137
S
Stanford University
Scholars:
9.6W
Papers: 8.2W
Citations: 17.0W