arrow
Return

Action-to-Action Diffusion Network for Weakly Supervised Temporal Action Localization

delete2025-01-01
delete0
PRE
AI
Y
Yuanbing Zou
赵清杰 (Qingjie Zhao)
P
Prodip Kumar Sarker
杨乐 (Le Yang)
B
Binglu Wang
DOI:10.1109/TMM.2025.3613174delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Weakly supervised temporal action localization (WTAL) aims to identify action instances in untrimmed videos with only video-level supervision. Despite recent advances in WTAL methods, achieving accurate boundary localization remains a significant challenge. A key reason is that WTAL networks following a localization-by-classification pipeline tend to focus on the most discriminative features, neglecting some ambiguous features that may contain action instances. To make the WTAL model focus on low-discriminative features that include action instances, we propose an action-to-action diffusion (ActionDiff) network. This network leverages the smoothness of data generated by the diffusion model, using the diffusion model to output smooth and high-quality features that weaken the discriminative action features from the base branch, thereby enhancing the performance of the WTAL task. First, we develop a topk-based masking strategy to generate binary masks that serve as pseudo-labels for diffusion model learning. Then, we propose a diffusion branch to generate high-quality latent action space by iteratively removing noise guided by the designed pseudo-labels and conditional information. To enhance the diffusion branch’s capability to generate human behavioral features, we design an action-related conditional strategy to obtain conditional information and use it to guide the modeling of human behavior knowledge by the diffusion branch. Our comprehensive experiments demonstrate that the proposed method achieves a promising performance on three benchmark datasets: THUMOS14, ActivityNet v1.2, and v1.3.
Keywords:
Temporal action localization
diffusion models
weakly supervised learning
temporal enhancement

Journal

IEEE Transactions on Multimedia cover
IEEE Transactions on Multimedia
IF:
9.7
Papers:
4.5K
Citations:
2.4W

Organization

H
Hefei Comprehensive National Science Center
Scholars:
200
Papers: 109
Citations: 0
B
beijing institute of technology
Scholars:
5.5W
Papers: 4.0W
Citations: 63
B
Begum Rokeya University
Scholars:
103
Papers: 78
Citations: 0
N
Northwestern Polytechnical University
Scholars:
4.6W
Papers: 3.7W
Citations: 5.3W
researcher View more organizations