arrow
Return

Monocular-Based 3-D Human Pose Estimation With Refinement Block and Special Loss Function

delete2025-02-01
delete0
delete
OA
AI
T
Tsung‐Han Tsai *
Y
Yi-Jhen Luo
DOI:10.1109/JSEN.2024.3510728delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Recently, 3-D human pose estimation (HPE) from a monocular RGB image has attracted much attention following the success of a deep convolution neural network (CNN). Many algorithms take 2.5-D heatmaps as the 3-D coordinate, whose X- and Y-axes correspond to the image coordinate, and the Z-axis corresponds to the camera coordinate. Therefore, the camera matrix or the distance between the root skeleton and the camera (the ground-truth information) is usually adopted to transform the 2.5-D coordinate to 3-D space. Since 2.5-D heatmaps ignore the conversion between 2-D and 3-D positions, they lose some conversion features and limit their applicability in the real world. In this article, we present an end-to-end framework that can utilize the contextual information in RGB images to directly predict 3-D space skeletons from a monocular image. Specifically, we use the multiloss method that depends on 2-D heatmaps, volumetric heatmaps, and a refinement block to locate the root-relative 3-D human pose. Our approach takes 2-D heatmaps and volumetric heatmaps as features to compute the loss and combine the loss from relative 3-D locations to generate the total loss. The model can learn the 2-D heatmap feature and 3-D location jointly and focus on the root-relative 3-D position in the camera coordinate. The experimental result shows that our model can predict relative 3-D human pose well on Human3.6M.
Keywords:
3-D human pose estimation (HPE)
deep convolution neural network (CNN)
root-relative 3-D human pose
volumetric heatmap

Journal

IEEE Sensors Journal cover
IEEE Sensors Journal
IF:
4.5
Papers:
2.1W
Citations:
7.3W

Organization

N
National Central University
Scholars:
1.0W
Papers: 8.5K
Citations: 6.4K