arrow
返回

Improving Weakly Supervised Temporal Action Localization by Exploiting Multi-Resolution Information in Temporal Domain

delete2021-01-01
delete12
PRE
AI
苏
苏瑞 (Rui Su)
Xu, Dong 封面图
Xu, Dong (Dong Xu) *
L
Luping Zhou
W
Wanli Ouyang
DOI:10.1109/TIP.2021.3089355delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Weakly supervised temporal action localization is a challenging task as only the video-level annotation is available during the training process. To address this problem, we propose a two-stage approach to generate high-quality frame-level pseudo labels by fully exploiting multi-resolution information in the temporal domain and complementary information between the appearance (i.e., RGB) and motion (i.e., optical flow) streams. In the first stage, we propose an Initial Label Generation (ILG) module to generate reliable initial frame-level pseudo labels. Specifically, in this newly proposed module, we exploit temporal multi-resolution consistency and cross-stream consistency to generate high quality class activation sequences (CASs), which consist of a number of sequences with each sequence measuring how likely each video frame belongs to one specific action class. In the second stage, we propose a Progressive Temporal Label Refinement (PTLR) framework to iteratively refine the pseudo labels, in which we use a set of selected frames with highly confident pseudo labels to progressively train two networks and better predict action class scores at each frame. Specifically, in our newly proposed PTLR framework, two networks called Network-OTS and Network-RTS, which are respectively used to generate CASs for the original temporal scale and the reduced temporal scales, are used as two streams (i.e., the OTS stream and the RTS stream) to refine the pseudo labels in turn. By this way, multi-resolution information in the temporal domain is exchanged at the pseudo label level, and our work can help improve each network/stream by exploiting the refined pseudo labels from another network/stream. Comprehensive experiments on two benchmark datasets THUMOS14 and ActivityNet v1.3 demonstrate the effectiveness of our newly proposed method for weakly supervised temporal action localization.
Keyword:
Videos
Location awareness
Task analysis
Feature extraction
Reliability
Annotations
Three-dimensional displays
Weakly supervised temporal action localization
temporal multi-resolution information
two stream fusion
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Image Processing 封面图
IEEE Transactions on Image Processing
IF:
13.7
论文数:
1.0W
被引数:
8.4W

机构

U
University of Sydney
学者数:
6.5W
论文数: 6.2W
被引数: 90
引用论文

引用论文

Spatial Enhancement and Temporal Constraint for Weakly Supervised Action Localization
err2020-01-01
err5
PREAI
errQin, Xiaolei; Ge, Yongxin; Yu, Hui; Chen, Feiyu; Yang, Dan
err分享
err收藏
err分享
err收藏
err分享
err收藏
err分享
err收藏
A new general method for protection of the hydroxyl function
err1976-03-01
err0
PREAI
errE.J. Corey; Jean-Louis Gras; Peter Ulrich
err分享
err收藏
A Methanol-Tolerant Gas-Venting Microchannel for a Microdirect Methanol Fuel Cell
err2007-12-01
err0
PREAI
errDennis Desheng Meng; Thomas Cubaud; Chih-Ming Ho; Chang-Jin Kim
err分享
err收藏
学者 查看更多内容