Return
Three-dimensional human pose estimation based on multi-scale spatial–temporal transformer
DOI:10.1016/j.engappai.2025.112068.png)
Abstract
En 中文
Recently, transformer-based methods have become dominant in the domain of three-dimensional (3D) human pose estimation, yet the U-net model based on convolutional neural networks (CNN) struggles to model long temporal sequences, and sequence-to-frame (Seq2frame) and sequence-to-sequence (Seq2seq) approaches often fail to preserve dependencies at the start and end of sequences. To address these challenges, this paper proposes a Multi-Scale Spatial–Temporal Transformer network (MSST). This network utilizes Sequence Padding Module (SPM) to extract edge features of the first and last frames, and employs Spatial–Temporal Transformer (STT) to model the spatial–temporal correlations of keypoints. Additionally, we design a Multi-Scale Module (MSM) that analyzes the human skeletal topology to extract multi-scale features of keypoints, local information, and global information, and fuse semantic information at different scales. Finally, we utilize regression heads to project the processed keypoint feature information into 3D space. We conduct quantitative evaluations on two benchmark datasets using four evaluation metrics and design multiple sets of comparative experiments to validate the effectiveness of the proposed modules. Experimental results demonstrate that the proposed network achieves excellent performance.
Keywords:
3D human pose estimation
transformer
convolutional neural networks
spatial-temporal modeling
multi-scale features
Journal
IF:
8
Papers:
5.3K
Citations:
3.5W

