返回
DDPTA: Zero-Shot Learning for Skeleton-Based Action Recognition
DOI:10.1109/LSP.2025.3650464.png)
摘要
En 中文
Traditional skeleton-based action recognition methods rely on large labeled datasets, which are costly to collect and unsuitable for hazardous actions, thereby limiting generalization. To overcome these limitations, recent works adopt zero-shot learning by using rich textual descriptions to guide the alignment and recognition of unlabeled skeleton features. However, these methods still struggle with similar actions (e.g., reading vs. writing), due to ambiguity arising from noise in both modalities. We propose the Discriminative Dual-Prototype TextAlignment (DDPTA) framework. Our framework introduces a novel dual-prototype design with tailored refinement strategies to effectively distill these two complementary prototypes. For the Spatial Prototype, our CycleSpatial module first distills the action’s core joint form from noisy spatial features, which is then guided by a Sieve-based Alignment. For the Temporal Prototype, our MambaTempo module leverages the Selective State Space Model to extract representations across distinct temporal stages, enabling fine-grained alignment with descriptions of different time periods. Extensive experiments demonstrate the superior performance of our method, showcasing its effectiveness in advancing the field of zero-shot skeleton-based action recognition.
Keyword:
Skeleton-based recognition
zero-shot
prototype
Mamba
sieve
期刊
I
IF:
3.9
论文数:
784
被引数:
0
机构
引用论文
Part-Aware Unified Representation of Language and Skeleton for Zero-Shot Action Recognition语言与骨骼的部件感知统一表征用于零样本动作识别

