Return
Exploring Supervised Contrastive Learning for Skeleton-Based Temporal Action Segmentation
DOI:10.1109/TCDS.2025.3532694.png)
Abstract
En 中文
Skeleton-based temporal action segmentation (STAS) takes long human skeleton sequences as input and predicts action categories at the frame level. Current STAS methods primarily use frame-wise cross-entropy loss, which focuses only on the relationships among one-hot label logits and overlooks the importance of the quality of frame-wise representations. To this end, we propose a novel framework called supervised contrastive skeleton-based temporal action segmentation (SCSAS) that optimizes representation learning. Specifically, our framework constructs a frame-level embedding space using a simple projection head and optimizes this space using three novel contrastive losses. These losses enhance the semantic relationships at the frame and segment levels by pulling together representations of the same activities and pushing apart those of different actions. Moreover, we introduce a confidence-based hard anchor sampling strategy to enhance the efficiency of contrastive learning. Finally, a boundary refinement module is also introduced to fully exploit the advantages of optimized representation. Our method seamlessly integrates with existing STAS methods and consistently enhances their performance across various datasets, without additional inference costs. With the incorporation of the boundary refinement branch, early STAS methods even achieve performance comparable to the state-of-the-art methods.
Keywords:
Boundary refinement
memory bank
skeleton-based temporal action segmentation (SCSAS)
supervised contrastive learning
Journal
IF:
4.9
Papers:
1.0K
Citations:
3.5K
Organization
No organization information available

