arrow
Return

Multi-Modality Self-Distillation for Weakly Supervised Temporal Action Localization

delete2022-01-01
delete16
PRE
AI
L
Linjiang Huang
W
Wang, Liang
H
Hongsheng Li *
DOI:10.1109/TIP.2021.3137649delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
As a challenging task of high-level video understanding, Weakly-supervised Temporal Action Localization (WTAL) has attracted increasing attention in recent years. However, due to the weak supervisions of whole-video classification labels, it is challenging to accurately determine action instance boundaries. To address this issue, pseudo-label-based methods [Alwassel et al. (2019), Luo et al. (2020), and Zhai et al. (2020)] were proposed to generate snippet-level pseudo labels from classification results. In spite of the promising performance, these methods hardly take full advantages of multiple modalities, i.e., RGB and optical flow sequences, to generate high quality pseudo labels. Most of them ignored how to mitigate the label noise, which hinders the capability of the network on learning discriminative feature representations. To address these challenges, we propose a Multi-Modality Self-Distillation (MMSD) framework, which contains two single-modal streams and a fused-modal stream to perform multi-modality knowledge distillation and multi-modality self-voting. On the one hand, multi-modality knowledge distillation improves snippet-level classification performance by transferring knowledge between single-modal streams and a fused-modal stream. On the other hand, multi-modality self-voting mitigates the label noise in a modality voting manner according to the reliability and complementarity of the streams. Experimental results on THUMOS14 and ActivityNet1.3 datasets demonstrate the effectiveness of our method and superior performance over state-of-the-art approaches. Our code is available at https://github.com/LeonHLJ/MMSD.
Keywords:
Location awareness
Reliability
Noise measurement
Annotations
Training
Head
Task analysis
Weakly supervised temporal action localization
multi-modality
pseudo label
self-distillation

Journal

IEEE Transactions on Image Processing cover
IEEE Transactions on Image Processing
IF:
13.7
Papers:
1.0W
Citations:
8.4W

Organization

I
institute of automation, cas
Scholars:
2.2K
Papers: 2.1K
Citations: 2
C
Chinese University of Hong Kong
Scholars:
3.4W
Papers: 3.2W
Citations: 5.6W
Cited Papers

Cited Papers

Relationship Between Amyloid β Protein and Melatonin Metabolite in a Study of Electric Utility Workers
err2002-08-01
err0
PREAI
errCurtis W. Noonan; John S. Reif; James B. Burch; Travers Y. Ichinose; Michael G. Yost; Kathy Magnusson
errShare
errSave
Timekeeping in genetically programmed aging
err1993-03-01
err0
PREAI
errP.E. Kloeden; R. Rössler; O.E. Rössler
errShare
errSave
Effect of Mn and C on Grain Growth in Mn Steels
err2018-11-29
err0
PREAI
errMadhumanti Bhattacharyya; Brian Langelier; Gary R. Purdy; Hatem S. Zurob
errShare
errSave
Duplex n- and p-Type Chromia Grown on Pure Chromium: A Photoelectrochemical and Microscopic Study
err2016-09-02
err0
PREAI
errL. Latu-Romain; Y. Parsa; S. Mathieu; M. Vilasi; M. Ollivier; A. Galerie; Y. Wouters
errShare
errSave
Modeling Sub-Actions for Weakly Supervised Temporal Action Localization
err2021-01-01
err24
PREAI
errHuang, Linjiang; Huang, Yan; Ouyang, Wanli; Wang, Liang
errShare
errSave
PREGNANCY HEPATITIS IN LIBYA
err1976-10-01
err0
PREAI
errA CHRISTIE
errShare
errSave
researcher View more