返回
Audio-visual self-supervised representation learning: A survey
DOI:10.1016/j.neucom.2025.129750.png)
摘要
En 中文
Artificial intelligence developers leverage the inherent relationships among video, text, and audio to create enhanced representations of the world, mirroring the way humans use multiple senses to understand their environment. As such, multimodal learning, which integrates various data input modalities to augment the learning of intrinsic features, has been gaining traction. While applications in multimodal understanding have made strides with deep learning, they often rely heavily on supervised learning and extensive human annotation. This paper provides a comprehensive review of audio-visual self-supervised learning, a promising alternative that uses vast amounts of unlabeled data. It holds the potential to reshape areas like computer vision, and speech recognition. We begin by explaining the concept of audio-visual modalities in machine learning and then move into their role within self-supervised learning by discussing terminology, general pipelines, and underlying motivations. This is followed by an exploration of common pretext tasks in audio- visual self-supervised learning, along with the evaluation methods, datasets, and downstream tasks associated with it. We then highlight prevailing challenges in both audio-visual and self-supervised learning realms. The paper concludes by presenting open challenges, suggesting avenues for future research in this dynamic domain.
Keyword:
Multimodal
Self-supervised learning
Deep learning
Pretext tasks
Data representation
Audio-visual learning
期刊
IF:
6.5
论文数:
2.5W
被引数:
6.5W

