arrow
返回

Audio-visual self-supervised representation learning: A survey

delete2025-02-01
delete0
PRE
AI
M
Manal S. Alsuwat *
S
Sarah Al-Shareef
M
Manal Alghamdi
DOI:10.1016/j.neucom.2025.129750delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Artificial intelligence developers leverage the inherent relationships among video, text, and audio to create enhanced representations of the world, mirroring the way humans use multiple senses to understand their environment. As such, multimodal learning, which integrates various data input modalities to augment the learning of intrinsic features, has been gaining traction. While applications in multimodal understanding have made strides with deep learning, they often rely heavily on supervised learning and extensive human annotation. This paper provides a comprehensive review of audio-visual self-supervised learning, a promising alternative that uses vast amounts of unlabeled data. It holds the potential to reshape areas like computer vision, and speech recognition. We begin by explaining the concept of audio-visual modalities in machine learning and then move into their role within self-supervised learning by discussing terminology, general pipelines, and underlying motivations. This is followed by an exploration of common pretext tasks in audio- visual self-supervised learning, along with the evaluation methods, datasets, and downstream tasks associated with it. We then highlight prevailing challenges in both audio-visual and self-supervised learning realms. The paper concludes by presenting open challenges, suggesting avenues for future research in this dynamic domain.
Keyword:
Multimodal
Self-supervised learning
Deep learning
Pretext tasks
Data representation
Audio-visual learning

期刊

Neurocomputing 封面图
Neurocomputing
IF:
6.5
论文数:
2.5W
被引数:
6.5W

机构

U
Umm Al Qura University
学者数:
2.9K
论文数: 3.1K
被引数: 4
引用论文

引用论文

Characterization of a shortened model of diet alternation in female rats
err2014-10-01
err0
errOAAI
errAngelo Blasio; Kenner C. Rice; Valentina Sabino; Pietro Cottone
err分享
err收藏
S-wave velocity structure and site effect parameters derived from microtremor arrays in the Western Plain of Taiwan
err2016-10-01
err0
errOAAI
errChun-Hsiang Kuo; Chun-Te Chen; Che-Min Lin; Kuo-Liang Wen; Jyun-Yan Huang; Shun-Chiang Chang
err分享
err收藏
err分享
err收藏
You Said That?: Synthesising Talking Faces from Audio
err2019-02-13
err101
errOAAI
errJamaludin, Amir; Chung, Joon Son; Zisserman, Andrew
err分享
err收藏
Sprint Ability: How Well Does Your Software Exploit Bursts in Processing Capacity?
err2016-07-01
err0
PREAI
errNathaniel Morris; Siva Meenakshi Renganathan; Christopher Stewart; Robert Birke; Lydia Chen
err分享
err收藏
Décollement séreux maculaire révélateur d’une leucémie aiguë lymphoblastique
err2005-01-01
err0
PREAI
errE. Abdallah; Z. Hajji; Z. Mellal; M. Belmekki; F. Bencherifa; A. Berraho
err分享
err收藏
err分享
err收藏
学者 查看更多内容