Return
Efficient 3D human pose estimation for IoT-based motion capture using Spatiotemporal Attention
DOI:10.1016/j.aej.2025.05.067.png)
Abstract
En 中文
With the growing demand for efficient and accurate 3D human pose estimation in fields such as virtual reality, human–computer interaction, sports analysis, and IoT-based monitoring, current Transformer-based solutions face challenges due to their quadratic computational cost as the number of joints and frames increases. To address this, we propose a 3D pose estimation network that combines Spatio-Temporal Criss-Cross Attention (STC) and a central point attention mechanism. The STC module splits the input features into spatial and temporal parts, applying self-attention to capture joint relationships within spatial frames and track dependencies across temporal frames. The central point attention mechanism uses a voxel network to refine pose regression within the central point range. By stacking multiple STC modules and introducing structure-enhanced positional embedding (SPE), our method captures spatiotemporal features and local structures. Experiments on the Human3.6M and MPI-INF-3DHP datasets show our approach achieves state-of-the-art accuracy with low computational cost, making it ideal for IoT-based monitoring and real-world applications requiring efficient pose estimation.
Keywords:
3D human pose estimation
Spatio-Temporal Criss-Cross Attention (STC)
Structure-enhanced Positional Embedding (SPE)
Computational efficiency
Internet of Things
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
6.8
Papers:
6.3K
Citations:
2.6W

