arrow
Return

Self-supervised 3D human pose estimation from video

delete2022-06-01
delete13
PRE
AI
M
Mohsen Gholami *
A
Ahmad Rezaei
H
Helge Rhodin
R
Rabab Ward
Z
Z. Jane Wang
DOI:10.1016/j.neucom.2022.02.076delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
To accurately estimate 3D human pose from monocular camera images, a large amount of 3D annotated data is required. However, obtaining 3D annotated data outside the laboratory is not easy. In the absence of such data, weakly-supervised methods that rely on multi-view cameras during training and single-view cameras during inference have been proposed. These methods either use multi-view networks or classical triangulation to train the 3D human pose estimator. This study shows that these two paradigms can collaborate to further improve performance. The available unlabeled uncalibrated multi-view inputs are used to obtain pseudo-3D labels employing classical triangulation. A pose estimator is trained with these pseudo-3D labels and with multi-view re-projection loss. This loss enforces the 3D poses estimated from different views to be consistent and improves the performance. Therefore, our method relaxes the constraints (calibrated cameras, 2D/3D annotations), only requires multi-view videos for training, and is therefore convenient for in-the-wild settings. The proposed method outperforms previous works on two challenging datasets, Human3.6 M and MPI-INF-3DHP. Codes and pretrained models will be publicly available.(c) 2022 Elsevier B.V. All rights reserved.
Keywords:
3D human pose
Self-supervised learning
Multi-view geometry
Triangulation

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

U
University of British Columbia
Scholars:
7.0W
Papers: 6.1W
Citations: 8.6W