返回
Adaptive 3D Human Pose Estimation Based on Spatial–Temporal Complexity Awareness
DOI:10.3390/electronics15102076.png)
摘要
En 中文
现有3D人体姿态估计算法采用固定的计算策略处理多样化的动作序列,导致简单动作计算冗余、复杂动作高频信息捕捉不足以及长序列处理效率低下。为解决这些问题,本文提出了一种空间-时序复杂度感知自适应计算框架(CAAPoseFormer)。首先,构建了空间-时序耦合复杂度量化模块,整合空间离散度和时序运动方差以实现分级动作复杂度量化。在此基础上,提出了时域-频域双域自适应剪枝策略,按需动态分配时序窗口长度和频域DCT系数。此外,设计了掩码引导的稀疏交互编码机制,通过屏蔽无效填充区域,实现变长特征的高效并行计算。在Human3.6M数据集上的实验表明,相较于基线PoseFormerV2,所提方法将参数量削减85.3%,计算成本降低64.8%,同时保持相当精度(MPJPE 44.2 mm),提升单元计算效率2.8倍。此外,与MHFormer和MotionBERT等当前最优(SOTA)方法相比,本方法将计算成本(MACs)分别降低97.4%和近三个数量级。该框架有效突破了高精度模型在低功耗硬件上的推理瓶颈,非常适合对延迟敏感的实时应用。
Keyword:
3D human pose estimation
transformer
spatial–temporal complexity-aware
time–frequency dual-domain adaptive pruning
mask-guided sparse interaction encoding
期刊
IF:
2.6
论文数:
1.0W
被引数:
4.7W
机构
引用论文
Chen, P.; Zeng, X.; Zhao, M.; Ye, P.; Shen, M.; Cheng, W.; Yu, G.; Chen, T. Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers. arXiv 2025, arXiv:2506.03065. [Google Scholar] [CrossRef]陈鹏; 曾晓; 赵敏; 叶平; 沈明; 程伟; 俞刚; 陈涛. Sparse-vDiT: 利用稀疏注意力加速视频扩散变换器. arXiv 2025, arXiv:2506.03065. [Google Scholar] [CrossRef]
Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition用于基于骨架的动作识别的时空图卷积网络
Shen, L.; Hao, T.; He, T.; Zhao, S.; Zhang, Y.; Liu, P.; Bao, Y.; Ding, G. TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval. In Proceedings of the 13th International Conference on Learning Representations (ICLR), Singapore, 24–28 April 2025. [Google Scholar] [CrossRef]沈, L.; 郝, T.; 何, T.; 赵, S.; 张, Y.; 刘, P.; 饶, Y.; 丁, G. TempMe: 用于高效文本-视频检索的视频时序标记合并. 在第13届学习表征国际会议(ICLR), 新加坡, 2025年4月24–28日. [Google Scholar] [CrossRef]
Wang, Y.; Chen, Z.; Jiang, H.; Li, S. Adaptive Computation Routing for Highly Efficient 3D Human Pose Estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; pp. 8812–8821. [Google Scholar] [CrossRef]王宇; 陈智; 姜浩; 李思. 高效3D人体姿态估计的自适应计算路由. 在IEEE/CVF国际计算机视觉会议(ICCV), 巴黎, 法国, 2023年10月1-6日; 第8812-8821页. [Google Scholar] [CrossRef]
Exploiting Temporal Contexts With Strided Transformer for 3D Human Pose Estimation利用带跨步变换的时间上下文进行3D人体姿态估计
Rao, Y.; Zhao, W.; Liu, B.; Lu, J.; Zhou, J.; Hsieh, C.-J. DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification. In Advances in Neural Information Processing Systems 34; Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., Vaughan, J.W., Eds.; Curran Associates, Inc.: Red Hook, NY, USA, 2021; pp. 13937–13949. [Google Scholar] [CrossRef]Rao, Y.; Zhao, W.; Liu, B.; Lu, J.; Zhou, J.; Hsieh, C.-J. DynamicViT:具有动态标记稀疏化的高效视觉变换器。载于《神经信息处理系统进展》第34卷;Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., Vaughan, J.W. 主编;Curran Associates, Inc.:纽约州红钩市,美国,2021年,第13937–13949页。[Google Scholar][CrossRef]
Liang, Y.; Ge, C.; Tong, Z.; Song, Y.; Wang, J.; Xie, P. Not All Patches Are What You Need: Expediting Vision Transformers via Token Reorganizations. In Proceedings of the 10th International Conference on Learning Representations (ICLR), Virtual, 25–29 April 2022. [Google Scholar] [CrossRef]梁, Y.; 葛, C.; 唐Z.; 宋, Y.; 王, J.; 谢P. 并非所有补丁都是你需要的:通过标记重组加速视觉变压器. 第10届学习表征国际会议(ICLR)会议录, 虚拟会议, 2022年4月25–29日. [Google Scholar] [CrossRef]

