返回
Comparing Learning Methodologies for Self-Supervised Audio-Visual Representation Learning
DOI:10.1109/ACCESS.2022.3164745.png)
摘要
En 中文
In recent years, the machine learning community has devoted an increasing attention to self-supervised learning.The performance gap between supervised and self-supervised has become increasingly narrow in many computer vision applications. In this paper, a new self-supervised approach is proposed for learning audio-visual representations from large databases of unlabeled videos. Our approach learns its representations by a combination of two tasks: unimodal and cross-modal. It uses a future prediction task, and learns to align its visual representations with its corresponding audio representations. To implement these tasks, three methodologies are assessed: contrastive learning, prototypical constrasting and redundancy reduction. The proposed approach is evaluated on a new publicly available dataset of videos captured from video game gameplay footage, called Videogame DB. On most downstream tasks, our method significantly outperforms baselines, demonstrating the real benefits of self-supervised learning in a real-world application.
Keyword:
Videos
Task analysis
Visualization
Feature extraction
Representation learning
Image recognition
Semantics
Self-supervised learning
audiovisual correspondence
cross-modal video representation learning
future prediction
learning methodologies
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
引用论文
Do Perceptions of Competence Mediate The Relationship Between Fundamental Motor Skill Proficiency and Physical Activity Levels of Children in Kindergarten?能力的感知是否可以介导幼儿园儿童的基本运动技能熟练程度与身体活动水平之间的关系?
Cobalt and Copper Composite Oxides as Efficient Catalysts for Preferential Oxidation of CO in H2-Rich Stream钴和铜复合氧化物作为H2-Rich流中CO优先氧化的有效催化剂

