arrow
Return

Snippet-to-Prototype Contrastive Consensus Network for Weakly Supervised Temporal Action Localization

delete2024-01-01
delete0
PRE
AI
Y
Yuxiang Shao
张飞飞 cover
张飞飞 (Feifei Zhang)
徐常胜 (Changsheng Xu) *
DOI:10.1109/TMM.2024.3355628delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Weakly-supervised temporal action localization aims to localize action instances from untrimmed videos with only video-level labels. Due to the lack of frame-wise annotations, most methods embrace a localization-by-classification paradigm. However, the large supervision gap between classification and localization hinders models from obtaining accurate snippet-wise classification sequences and action proposals. We propose a snippet-to-prototype contrastive consensus network (SPCC-Net) to simultaneously generate feature-level and label-level supervision information to narrow the supervision gap between classification and localization. Specifically, the network adopts a two-stream framework incorporating the optical flow and fusion streams to fully leverage the motion and complementary information from multiple modalities. Firstly, the snippet-to-prototype contrast module is executed within each stream to learn prototypes for all categories and contrast them with action snippets to guarantee intra-class compactness and inter-class separability of snippet features. Secondly, for generating accurate label-level supervision information through complementary information of multimodal features, the multi-modality consensus module ensures not only category consistency through knowledge distillation but also semantic consistency through contrastive learning. Finally, we introduce the auxiliary multiple instance learning (MIL) loss to alleviate the issue that existing MIL-based methods only localize sparse discriminative snippets. Extensive experiments are conducted on two public datasets, THUMOS-14 and ActivityNet-1.3, to demonstrate the superior performance of our method over state-of-the-art methods.
Keywords:
Contrastive learning
knowledge distillation
weakly-supervised temporal action localization

Journal

IEEE Transactions on Multimedia cover
IEEE Transactions on Multimedia
IF:
9.7
Papers:
4.5K
Citations:
2.4W

Organization

T
Tianjin University of Technology
Scholars:
8.8K
Papers: 5.9K
Citations: 1.0W
C
chinese academy of sciences
Scholars:
56.4W
Papers: 44.9W
Citations: 704