arrow
Return

Exploring multi-level transformers with feature frame padding network for 3D human pose estimation

delete2024-08-13
delete2
PRE
AI
J
Jae Hoon Jeong
Y
Young Hoon Joo *
DOI:10.1007/s00530-024-01451-4delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Recently, transformer-based architecture achieved remarkable performance in 2D to 3D lifting pose estimation. Despite advancements in transformer-based architecture they still struggle to handle depth ambiguity, limited temporal information, lacking edge frame details, and short-term temporal features. Consequently, transformer architecture encounters challenges in preciously estimating the 3D human position. To address these problems, we proposed Multi-Level Transformers with a Feature Frame Padding Network (MLTFFPN). To do this, we first propose the frame-padding network, which allows the network to capture longer temporal dependencies and effectively address the lacking edge frame information, enabling a better understanding of the sequential nature of human motion and improving the accuracy of pose estimation. Furthermore, we employ a multi-level transformer to extract temporal information from 3D human poses, which aims to improve the short-range temporal dependencies among keypoints of the human pose skeleton. Specifically, we introduce the Refined Temporal Constriction and Proliferation Transformer (RTCPT), which incorporates spatio-temporal encoders and a Temporal Constriction and Proliferation (TCP) structure to reveal multi-scale attention information and effectively addresses the depth ambiguity problem. Moreover, we incorporate the Feature Aggregation Refinement (FAR) module into the TCP block in a cross-layer manner, which facilitates semantic representation through the persistent interaction of queries, keys, and values. We extensively evaluate the efficiency of our method through experiments on two well-known benchmark datasets: Human3.6M and MPI-INF-3DHP.
Keywords:
3D human pose estimation
Frame-padding network
Spatio-temporal transformer
Temporal constriction and proliferation transformer
Feature aggregation refinement module

Journal

Multimedia Systems cover
Multimedia Systems
IF:
3.1
Papers:
2.8K
Citations:
2.7K

Organization

K
Kunsan National University
Scholars:
1.5K
Papers: 1.7K
Citations: 1.4K
Cited Papers

Cited Papers

Exploiting Temporal Contexts With Strided Transformer for 3D Human Pose Estimation
err2023-01-01
err84
errOAAI
errLi, Wenhao; Liu, Hong; Ding, Runwei; Liu, Mengyuan; Wang, Pichao; Yang, Wenming
errShare
errSave
MicroRNA-23a expression in paraffin-embedded specimen correlates with overall survival of diffuse large B-cell lymphoma
err2014-03-22
err0
PREAI
errWan-ling Wang; Cui Yang; Xiao-lin Han; Ran Wang; Yan Huang; You-mei Zi; Jing-dong Li
errShare
errSave
err
IF0
err
err0
PREAI
err
errShare
errSave
errShare
errSave
Nanoceramic and Polytetrafluoroethylene Polymer Composites for Mechanical Seal Application at Low Temperature
err2013-05-20
err0
errOAAI
errA.A. Okhlopkova; S.A. Sleptsova; G.N. Alexandrov; A.E. Dedyukin; Ee Le Shim; Dae-Yong Jeong; Jin-Ho Cho
errShare
errSave
researcher View more