Return
Spatio-temporal feature discrimination for self-supervised skeleton action representation learning
DOI:10.1007/s00530-026-02421-8.png)
Abstract
En 中文
This paper proposes a self-supervised skeleton learning framework STFD based on spatio-temporal feature discrimination for skeleton sequence representation learning. Existing unsupervised methods typically rely on global data augmentations and overlook local spatio-temporal relations among joints in skeleton sequences. To address this, STFD introduces spatio-temporal masking strategies that mask skeleton sequences guided by local relations, improving recognition accuracy and robustness. Concretely, the spatio-temporal part adopts a Trunk Spatial Mask (TSM) and Selective Temporal Suppression (STS) to capture local structures across both spatial and temporal dimensions, respectively. In the architecture for feature discrimination, This paper perform feature decoupling and frame-level feature discrimination on skeleton data, and jointly discriminate the representations across the joint,bone,motion three-stream inputs. Experiments show that STFD achieves performance comparable to the leading methods on the NTU RGB+D 60, NTU RGB+D 120, and PKU-MMD datasets.Moreover, STFD remains highly robust when parts of skeleton joints are missing, validating its effectiveness and broad applicability under incomplete-data scenarios.
Keywords:
Action recognition
Spatio-temporal feature discrimination
Spatio-temporal feature decoupling
Self-supervised learning
Journal
IF:
3.1
Papers:
2.7K
Citations:
2.7K

