arrow
Return

Vision-language model guided pose knowledge mining for human pose estimation

delete2025-09-01
delete0
delete
OA
AI
C
Chen, Yilei
X
Xuemei Xie *
F
Fu Li *
DOI:10.1093/jcde/qwaf079delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Vision-language models with large-scale image-text pairs have shown significant potential on representation learning. Human pose estimation task, which is highly sensitive to pixel-wise transformation, requires effective methods for mining pose-specific knowledge. In this paper, we investigate the homologous human pose retrieval task relying on large-scale annotated datasets to enhance pose knowledge extraction. We propose Pose Prompt (PosePro), which leverages vision-language models to categorize global pose configuration of an image, build compatible design, generate pose embedding as proposals. We then aim to integrate the learned knowledge as visual and textual prompt to facilitate the learning processing of newly unseen tasks. We demonstrate the effectiveness of fundamental PosePro model through extensive experiments on both pose retrieval and human pose estimation, showing significant improvements in accuracy and generalization ability, especially in scenarios with limited samples.
Keywords:
vision-language model
human pose estimation
knowledge
transfer learning
generalization

Journal

Journal of Computational Design and Engineering cover
Journal of Computational Design and Engineering
IF:
6.1
Papers:
393
Citations:
3.2K

Organization

X
Xidian University
Scholars:
2.4W
Papers: 1.9W
Citations: 9.7K