Return
PSINet: Parallel Space Interaction Network for Human Pose Estimation
DOI:10.1109/TETCI.2025.3647687.png)
Abstract
En 中文
Human pose estimation is sensitive to the spatial resolution of feature representations in a single image. While multi-scale features are often exploited to improve keypoint localization, the effectiveness of parallel spatial learning across multiple uniform-sized spaces remains underexplored. To address this, To address this, we propose the Parallel Space Interaction Network (PSINet), which splits input features into multiple identical parallel spaces to facilitate efficient spatial feature learning for human pose estimation. Based on the multiple spaces, we refine the vanilla vision Transformer by proposing a Split-Head Transformer (SHT) to enhance global interaction within parallel spaces. Furthermore, we introduce Parallel Space Fusion instrument to strengthen the capacity of the PSINet by acquiring the relationship between multiple parallel spaces. Extensive experiments demonstrate that PSINet achieves state-of-the-art accuracy after data augmentation, running at 180 FPS on a single GTX2080Ti, offering a superior trade-off between accuracy and efficiency.
Keywords:
Parallel space interaction
transformer
human pose estimation
feature interaction
Journal
I
IF:
6.5
Papers:
1.4K
Citations:
4.5K

