Return
Multi-View Knowledge Guided Semantic Prototype Learning for Generalized Zero-Shot Action Recognition
DOI:10.1109/TMM.2025.3618570.png)
Abstract
En 中文
Generalized zero-shot skeleton-based action recognition (GZSSAR) is an emerging and challenging problem in the computer vision community. It requires models to recognize human actions, including some classes that are unseen during training. Previous studies typically rely solely on action labels to bridge the gap between seen and unseen action classes. However, the limited action semantic information hinders the learning of comprehensive semantic prototypes, thereby restricting the model’s ability to generalize to unseen classes. To address this issue, in addition to the original action labels, we explore four types of textual action descriptions (<italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">i.e</i>., interpretive and motional descriptions derived from manual expert annotation and large language model) for each action class. In order to comprehensively utilize multi-view semantic information for zero-shot classification, an Attentional Multi-view Semantic Fusion (AMSF) model is proposed. It effectively integrates the multi-view semantic features and aligns visual and semantic features in a common space, subsequently realizing the recognition of unseen action classes. Furthermore, previous works typically evaluate models in settings that include specific unseen classes, which is insufficient for GZSSAR research. To thoroughly evaluate different models, we introduce two novel distinct experimental settings, termed the “easy setting” and the “hard setting”, based on the semantic similarities between action classes. Extensive experimental results on three large-scale skeleton-based action recognition benchmarks (PKUMMD, NTU-60, and NTU-120) not only validate the advantages of the proposed multiview action descriptions and the AMSF model but also demonstrate the rationality of the novel experimental settings.
Keywords:
Skeleton-based action recognition
generalized zero-shot learning
action descriptions
semantic fusion
Journal
IF:
9.7
Papers:
4.5K
Citations:
2.4W

