返回
Listen, Denoise, Action! Audio-Driven Motion Synthesis with Diffusion Models
DOI:10.1145/3592458.png)
摘要
En 中文
Diffusion models have experienced a surge of interest as highly expressive yet efficiently trainable probabilistic models. We show that these models are an excellent fit for synthesising human motion that co-occurs with audio, e.g., dancing and co-speech gesticulation, since motion is complex and highly ambiguous given audio, calling for a probabilistic description. Specifically, we adapt the DiffWave architecture to model 3D pose sequences, putting Conformers in place of dilated convolutions for improved modelling power. We also demonstrate control over motion style, using classifier-free guidance to adjust the strength of the stylistic expression. Experiments on gesture and dance generation confirm that the proposed method achieves top-of-the-line motion quality, with distinctive styles whose expression can be made more or less pronounced. We also synthesise path-driven locomotion using the same model architecture. Finally, we generalise the guidance procedure to obtain product-of-expert ensembles of diffusion models and demonstrate how these may be used for, e.g., style interpolation, a contribution we believe is of independent interest.
Keyword:
Generative models
machine learning
diffusion models
conformers
gestures
dance
locomotion
product of experts
ensemble models
guided interpolation
期刊
IF:
9.5
论文数:
4.7K
被引数:
3.6W
机构
引用论文
Sodium channel kinetic changes that produce Brugada syndrome or progressive cardiac conduction system disease产生Brugada综合征或进行性心脏传导系统疾病的钠通道动力学变化
Structural neural correlates of prosaccade and antisaccade eye movements in healthy humans
NeuroImage
IF0
Sensorless Control of Z Source Inverter fed BLDC Motor Drive by FOC - DTC Hybrid Control Strategy Using Fuzzy Logic Controller采用模糊逻辑控制器的foc-dtc混合控制策略的Z源逆变器馈电BLDC电机驱动的无传感器控制

