arrow
Return

DDPTA: Zero-Shot Learning for Skeleton-Based Action Recognition

delete2026-01-01
delete0
PRE
AI
J
Jinjie Wang
B
Bi Zeng
S
Shenghong Zhong
P
Pengfei Wei
X
Xiaoting Gao
DOI:10.1109/LSP.2025.3650464delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Traditional skeleton-based action recognition methods rely on large labeled datasets, which are costly to collect and unsuitable for hazardous actions, thereby limiting generalization. To overcome these limitations, recent works adopt zero-shot learning by using rich textual descriptions to guide the alignment and recognition of unlabeled skeleton features. However, these methods still struggle with similar actions (e.g., reading vs. writing), due to ambiguity arising from noise in both modalities. We propose the Discriminative Dual-Prototype TextAlignment (DDPTA) framework. Our framework introduces a novel dual-prototype design with tailored refinement strategies to effectively distill these two complementary prototypes. For the Spatial Prototype, our CycleSpatial module first distills the action’s core joint form from noisy spatial features, which is then guided by a Sieve-based Alignment. For the Temporal Prototype, our MambaTempo module leverages the Selective State Space Model to extract representations across distinct temporal stages, enabling fine-grained alignment with descriptions of different time periods. Extensive experiments demonstrate the superior performance of our method, showcasing its effectiveness in advancing the field of zero-shot skeleton-based action recognition.
Keywords:
Skeleton-based recognition
zero-shot
prototype
Mamba
sieve

Journal

I
IEEE Signal Processing Letters
IF:
3.9
Papers:
610
Citations:
0

Organization

G
guangdong university of technology
Scholars:
2.9W
Papers: 2.0W
Citations: 36