arrow
Return

Comparing Learning Methodologies for Self-Supervised Audio-Visual Representation Learning

delete2022-01-01
delete8
delete
OA
AI
H
Hacene Terbouche
L
Liam Schoneveld
O
Oisin Benson
A
Alice Othmani *
DOI:10.1109/ACCESS.2022.3164745delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In recent years, the machine learning community has devoted an increasing attention to self-supervised learning.The performance gap between supervised and self-supervised has become increasingly narrow in many computer vision applications. In this paper, a new self-supervised approach is proposed for learning audio-visual representations from large databases of unlabeled videos. Our approach learns its representations by a combination of two tasks: unimodal and cross-modal. It uses a future prediction task, and learns to align its visual representations with its corresponding audio representations. To implement these tasks, three methodologies are assessed: contrastive learning, prototypical constrasting and redundancy reduction. The proposed approach is evaluated on a new publicly available dataset of videos captured from video game gameplay footage, called Videogame DB. On most downstream tasks, our method significantly outperforms baselines, demonstrating the real benefits of self-supervised learning in a real-world application.
Keywords:
Videos
Task analysis
Visualization
Feature extraction
Representation learning
Image recognition
Semantics
Self-supervised learning
audiovisual correspondence
cross-modal video representation learning
future prediction
learning methodologies

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

U
universite paris-est-creteil-val-de-marne (upec)
Scholars:
1.3W
Papers: 9.2K
Citations: 6