arrow
Return

Action-Semantic Consistent Knowledge for Weakly-Supervised Action Localization

delete2024-01-01
delete0
PRE
AI
汪昱 cover
汪昱 (Wang, Yu)
S
Shengjie Zhao *
S
Shiwei Chen
DOI:10.1109/TMM.2024.3405710delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Weakly-supervised temporal action localization aims to detect temporal intervals of actions in arbitrarily long untrimmed videos with only video-level annotations. Owing to label sparsity, learning action consistency is intractable. In this paper, we assume that frames with similar representations in a given video should be considered as the same action. To this end, we develop a query-based contrastive learning paradigm to ensure action-semantic consistency. This mechanism encourages normalized embeddings with the same class to be pulled closer together, while embeddings from different classes are repelled apart. Besides, we design a two-branch framework, consisting of a class-aware branch and a class-agnostic branch, to learn salient features and fine-grained clues respectively. To further guarantee the action-semantic consistency of the two branches, unlike previous methods that handle each branch independently, we model the relationship between the two branches to avoid unreasonable predictions. Finally, the proposed model demonstrates superior performance over existing methods on the publicly available THUMOS-14 and ActivityNet-1.3 datasets. Substantial experiments and ablation studies also demonstrate the effectiveness of our model.
Keywords:
Videos
Feature extraction
Location awareness
Semantics
Annotations
Task analysis
Proposals
Temporal action localization
contrastive learning
action-semantic consistency

Journal

IEEE Transactions on Multimedia cover
IEEE Transactions on Multimedia
IF:
9.7
Papers:
4.5K
Citations:
2.4W

Organization

T
tongji university
Scholars:
7.7W
Papers: 5.9W
Citations: 98