Return
Semantics-guided topology fusion graph convolutional network for efficient skeleton-based action recognition
J
H
X
K
T
B
B
M
DOI:10.1016/j.knosys.2026.116770.png)
Abstract
En 中文
Graph topology is crucial for skeleton-based action recognition, as it determines the propagation of motion information among human joints. Existing GCN-based methods typically rely on physical skeleton connections or adaptive topologies learned from joint coordinates. However, such topologies are mainly constrained by anatomical structures or motion statistics, causing difficulty in capturing high-level semantic dependencies between physically distant but action-relevant joints. This limitation is particularly evident in fine-grained actions, where discriminative cues often depend on body-part functions, action context, and class-specific joint interactions. To address this issue, we propose a semantics-guided topology fusion graph convolutional network (SGTF-GCN), a novel framework that exploits large language models (LLMs) to derive offline semantic graph-topology priors to enhance skeleton-based action recognition. SGTF-GCN introduces two complementary structured priors: a global action-context prior for encoding holistic action semantics and a local joint-relation prior for capturing fine-grained functional joint dependencies. Both priors conform to the standard GCN adjacency matrix format, enabling seamless fusion with the original physical skeleton topology. Subsequently, a dedicated semantics-guided topology fusion module dynamically integrates these semantic priors with the physical topology, resulting in a strong semantics-guided multimodal teacher network. To maintain efficient inference without requiring LLMs at test time, we introduce topological knowledge distillation, which transfers the teacher’s multi-level semantic structural knowledge to a lightweight unimodal student network by aligning their topology distributions. Extensive experimental results reveal that the proposed method achieves competitive performance on NTU RGB+D, NTU RGB+D 120, and NW-UCLA with excellent inference efficiency.
Keywords:
Skeleton-based action recognition
Graph convolutional network
Semantic topology
Knowledge distillation
Large language models
Journal
K
IF:
7.6
Papers:
1.2W
Citations:
4.5W
