arrow
返回

Multi-hypothesis representation learning for transformer-based 3D human pose estimation

delete2023-09-01
delete12
PRE
AI
W
Wenhao Li
刘宏 封面图
刘宏 (Hong Liu) *
H
Hao Tang
P
Pichao Wang
DOI:10.1016/j.patcog.2023.109631delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Despite significant progress, estimating 3D human poses from monocular videos remains a challenging task due to depth ambiguity and self-occlusion. Most existing works attempt to solve both issues by ex-ploiting spatial and temporal relationships. However, those works ignore the fact that it is an inverse problem where multiple feasible solutions (i.e., hypotheses) exist. To relieve this limitation, we propose a Multi-Hypothesis Transformer that learns spatio-temporal representations of multiple plausible pose hypotheses. In order to effectively model multi-hypothesis dependencies and build strong relationships across hypothesis features, we introduce a one-to-many-to-one three-stage framework: (i) Generate mul-tiple initial hypothesis representations; (ii) Model self-hypothesis communication, merge multiple hy-potheses into a single converged representation and then partition it into several diverged hypotheses; (iii) Learn cross-hypothesis communication and aggregate the multi-hypothesis features to synthesize the final 3D pose. Through the above processes, the final representation is enhanced and the synthesized pose is much more accurate. Extensive experiments show that the proposed method achieves state-of -the-art results on two challenging datasets: Human3.6M and MPI-INF-3DHP. The code and models are available at https://github.com/Vegetebird/MHFormer .(c) 2023 Elsevier Ltd. All rights reserved.
Keyword:
3D Human pose estimation
Transformer
Multi-Hypothesis
Self-Hypothesis
Cross-Hypothesis

期刊

Pattern Recognition 封面图
Pattern Recognition
IF:
7.6
论文数:
1.3W
被引数:
4.5W

机构

E
ETH Zurich
学者数:
3.0W
论文数: 2.4W
被引数: 8.4W
S
swiss federal institutes of technology domain
学者数:
9.0W
论文数: 8.0W
被引数: 163
P
peking university
学者数:
11.9W
论文数: 8.7W
被引数: 146
学者 查看更多机构
引用论文

引用论文

err
IF0
err
err0
PREAI
err
err分享
err收藏
Enhanced 3D Human Pose Estimation from Videos by Using Attention-Based Neural Network with Dilated Convolutions
err2021-02-26
err22
PREAI
errLiu, Ruixu; Shen, Ju; Wang, He; Chen, Chen; Cheung, Sen-ching; Asari, Vijayan K.
err分享
err收藏
Effects of Expressive Writing Effects on Disgust and Anxiety in a Subsequent Dissection
err2014-09-20
err0
PREAI
errChristoph Randler; Peter Wüst-Ackermann; Viola Otte im Kampe; Inga H. Meyer-Ahrens; Benjamin J. Tempel; Christian Vollmer
err分享
err收藏
VNect: Real-time 3D Human Pose Estimation with a Single RGB Camera
err2017-07-20
err798
errOAAI
errMehta, Dushyant; Sridhar, Srinath; Sotnychenko, Oleksandr; Rhodin, Helge; Shafiei, Mohammad; Seidel, Hans-Peter; Xu, Weipeng; Casas, Dan; Theobalt, Christian
err分享
err收藏
没有更多内容