Return
DDPTA: Zero-Shot Learning for Skeleton-Based Action Recognition
DOI:10.1109/LSP.2025.3650464.png)
Abstract
En 中文
Traditional skeleton-based action recognition methods rely on large labeled datasets, which are costly to collect and unsuitable for hazardous actions, thereby limiting generalization. To overcome these limitations, recent works adopt zero-shot learning by using rich textual descriptions to guide the alignment and recognition of unlabeled skeleton features. However, these methods still struggle with similar actions (e.g., reading vs. writing), due to ambiguity arising from noise in both modalities. We propose the Discriminative Dual-Prototype TextAlignment (DDPTA) framework. Our framework introduces a novel dual-prototype design with tailored refinement strategies to effectively distill these two complementary prototypes. For the Spatial Prototype, our CycleSpatial module first distills the action’s core joint form from noisy spatial features, which is then guided by a Sieve-based Alignment. For the Temporal Prototype, our MambaTempo module leverages the Selective State Space Model to extract representations across distinct temporal stages, enabling fine-grained alignment with descriptions of different time periods. Extensive experiments demonstrate the superior performance of our method, showcasing its effectiveness in advancing the field of zero-shot skeleton-based action recognition.
Keywords:
Skeleton-based recognition
zero-shot
prototype
Mamba
sieve
Journal
I
IF:
3.9
Papers:
610
Citations:
0

