arrow
Return

Multi-View Knowledge Guided Semantic Prototype Learning for Generalized Zero-Shot Action Recognition

delete2025-01-01
delete0
PRE
AI
M
Mingzhe Li
Z
Zhen Jia
张彰 (Zhang Zhang)
Y
Yaoning Li
马占宇 (Zhanyu Ma)
王亮 cover
王亮 (Liang Wang)
DOI:10.1109/TMM.2025.3618570delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Generalized zero-shot skeleton-based action recognition (GZSSAR) is an emerging and challenging problem in the computer vision community. It requires models to recognize human actions, including some classes that are unseen during training. Previous studies typically rely solely on action labels to bridge the gap between seen and unseen action classes. However, the limited action semantic information hinders the learning of comprehensive semantic prototypes, thereby restricting the model’s ability to generalize to unseen classes. To address this issue, in addition to the original action labels, we explore four types of textual action descriptions (<italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">i.e</i>., interpretive and motional descriptions derived from manual expert annotation and large language model) for each action class. In order to comprehensively utilize multi-view semantic information for zero-shot classification, an Attentional Multi-view Semantic Fusion (AMSF) model is proposed. It effectively integrates the multi-view semantic features and aligns visual and semantic features in a common space, subsequently realizing the recognition of unseen action classes. Furthermore, previous works typically evaluate models in settings that include specific unseen classes, which is insufficient for GZSSAR research. To thoroughly evaluate different models, we introduce two novel distinct experimental settings, termed the “easy setting” and the “hard setting”, based on the semantic similarities between action classes. Extensive experimental results on three large-scale skeleton-based action recognition benchmarks (PKUMMD, NTU-60, and NTU-120) not only validate the advantages of the proposed multiview action descriptions and the AMSF model but also demonstrate the rationality of the novel experimental settings.
Keywords:
Skeleton-based action recognition
generalized zero-shot learning
action descriptions
semantic fusion

Journal

IEEE Transactions on Multimedia cover
IEEE Transactions on Multimedia
IF:
9.7
Papers:
4.5K
Citations:
2.4W

Organization

I
Institute of Automation
Scholars:
529
Papers: 278
Citations: 220