arrow
返回

Text2Human: Text-Driven Controllable Human Image Generation

delete2022-07-22
delete53
PRE
AI
Y
Yuming Jiang
S
Shuai Yang
Q
Qju, Haonan
W
Wayne Wu
C
Chen Change Loy
Z
Ziwei Liu *
DOI:10.1145/3528223.3530104delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Generating high-quality and diverse human images is an important yet challenging task in vision and graphics. However, existing generative models often fall short under the high diversity of clothing shapes and textures. Furthermore, the generation process is even desired to be intuitively controllable for layman users. In this work, we present a text-driven controllable framework, Text2Human, for a high-quality and diverse human generation. We synthesize full-body human images starting from a given human pose with two dedicated steps. 1) With some texts describing the shapes of clothes, the given human pose is first translated to a human parsing map. 2) The final human image is then generated by providing the system with more attributes about the textures of clothes. Specifically, to model the diversity of clothing textures, we build a hierarchical texture-aware codebook that stores multi-scale neural representations for each type of texture. The codebook at the coarse level includes the structural representations of textures, while the codebook at the fine level focuses on the details of textures. To make use of the learned hierarchical codebook to synthesize desired images, a diffusion-based transformer sampler with mixture of experts is firstly employed to sample indices from the coarsest level of the codebook, which then is used to predict the indices of the codebook at finer levels. The predicted indices at different levels are translated to human images by the decoder learned accompanied with hierarchical codebooks. The use of mixture-of-experts allows for the generated image conditioned on the fine-grained text input. The prediction for finer level indices refines the quality of clothing textures. Extensive quantitative and qualitative evaluations demonstrate that our proposed Text2Human framework can generate more diverse and realistic human images compared to state-of-the-art methods.
Keyword:
Image generation
controllable human image generation
text-driven generation

期刊

ACM Transactions on Graphics 封面图
ACM Transactions on Graphics
IF:
9.5
论文数:
4.7K
被引数:
3.6W

机构

N
Nanyang Technological University
学者数:
4.9W
论文数: 4.8W
被引数: 8.1W
引用论文

引用论文

Ki-67 Proliferation Index
err2004-02-01
err0
PREAI
errMichael G. Alexandrakis; Freda H. Passam; Despina S. Kyriakou; Konstantina Dambaki; Maria Niniraki; Efstathios Stathopoulos
err分享
err收藏
Race differences in the association of spiritual experiences and life satisfaction in older age
err2013-09-01
err0
errOAAI
errKimberly A. Skarupski; George Fitchett; Denis A. Evans; Carlos F. Mendes de Leon
err分享
err收藏
Use of a soil moisture network for drought monitoring in the Czech Republic
err2011-06-09
err0
PREAI
errMartin Mozny; Mirek Trnka; Zdenek Zalud; Petr Hlavinka; Jiri Nekovar; Vera Potop; Michal Virag
err分享
err收藏
Pose with Style: Detail-Preserving Pose-Guided Image Synthesis with Conditional StyleGAN
err2021-12-10
err64
PREAI
errAlbahar, Badour; Lu, Jingwan; Yang, Jimei; Shu, Zhixin; Shechtman, Eli; Huang, Jia-Bin
err分享
err收藏
Gene expression in retinal ischemic post-conditioning
err2018-03-05
err0
errOAAI
errKonrad Kadzielawa; Biji Mathew; Clara R. Stelman; Arden Zhengdeng Lei; Leianne Torres; Steven Roth
err分享
err收藏
err分享
err收藏
学者 查看更多内容