arrow
Return

Audio-Visual Deep Clustering for Speech Separation

delete2019-11-01
delete36
delete
OA
AI
R
Rui Lü *
Z
Zhiyao Duan
C
Changshui Zhang
DOI:10.1109/TASLP.2019.2928140delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Speech separation aims to separate individual voices from an audio mixture of multiple simultaneous talkers. Audio-only approaches show unsatisfactory performance when the speakers are of the same gender or share similar voice characteristics. This is due to challenges on learning appropriate feature representations for separating voices in single frames and streaming voices across time. Visual signals of speech (e.g., lip movements), if available, can be leveraged to learn better feature representations for separation. In this paper, we propose a novel audio-visual deep clustering model (AVDC) to integrate visual information into the process of learning better feature representations (embeddings) for Time-Frequency (T-F) bin clustering. It employs a two-stage audio-visual fusion strategy where speaker-wise audio-visual T-F embeddings are first computed after the first-stage fusion to model the audio-visual correspondence for each speaker. In the second-stage fusion, audio-visual embeddings of all speakers and audio embeddings calculated by deep clustering from the audio mixture are concatenated to form the final T-F embedding for clustering. Through a series of experiments, the proposed AVDC model is shown to outperform the audio-only deep clustering and utterance-level permutation invariant training baselines and three other state-of-the-art audio-visual approaches. Further analyses show that the AVDC model learns a better T-F embedding for alleviating the source permutation problem across frames. Other experiments show that the AVDC model is able to generalize across different numbers of speakers between training and testing and shows some robustness when visual information is partially missing.
Keywords:
Speaker-independent speech separation
deep clustering
audio-visual fusion
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

I
IEEE-ACM Transactions on Audio Speech and Language Processing
IF:
5.1
Papers:
2.6K
Citations:
1.1W

Organization

T
tsinghua university
Scholars:
11.8W
Papers: 10.0W
Citations: 137
U
University of Rochester
Scholars:
2.6W
Papers: 2.1W
Citations: 2.2W