返回
Deep Audio-visual Learning: A Survey
DOI:10.1007/s11633-021-1293-0.png)
摘要
En 中文
Audio-visual learning, aimed at exploiting the relationship between audio and visual modalities, has drawn considerable attention since deep learning started to be used successfully. Researchers tend to leverage these two modalities to improve the performance of previously considered single-modality tasks or address new challenging problems. In this paper, we provide a comprehensive survey of recent audio-visual learning development. We divide the current audio-visual learning tasks into four different subfields: audio-visual separation and localization, audio-visual correspondence learning, audio-visual generation, and audio-visual representation learning. State-of-the-art methods, as well as the remaining challenges of each subfield, are further discussed. Finally, we summarize the commonly used datasets and challenges.
Keyword:
Deep audio-visual learning
audio-visual separation and localization
correspondence learning
generative models
representation learning
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
I
IF:
3.7
论文数:
144
被引数:
1.4K
机构
引用论文
Characterization of the Corrosion Products Formed on Carbon Steel in Qinghai Salt Lake Atmosphere
Corrosion
IF0

