arrow
返回

Multi-Modal Multi-Channel Target Speech Separation

delete2020-03-01
delete62
delete
OA
AI
R
Rongzhi Gu
S
Shixiong Zhang
Y
Yong Xu
Y
Yuexian Zou *
D
Dong Yu
DOI:10.1109/JSTSP.2020.2980956delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Target speech separation refers to extracting a target speaker's voice from an overlapped audio of simultaneous talkers. Previously the use of visual modality for target speech separation has demonstrated great potentials. This work proposes a general multi-modal framework for target speech separation by utilizing all the available information of the target speaker, including his/her spatial location, voice characteristics and lip movements. Also, under this framework, we investigate on the fusion methods for multi-modal joint modeling. A factorized attention-based fusion method is proposed to aggregate the high-level semantic information of multi-modalities at embedding level. This method firstly factorizes the mixture audio into a set of acoustic subspaces, then leverages the target's information from other modalities to enhance these subspace acoustic embeddings with a learnable attention scheme. To validate the robustness of proposed multi-modal separation model in practical scenarios, the system was evaluated under the condition that one of the modalities is temporarily missing, invalid or corrupted. Experiments are conducted on a large-scale audio-visual dataset collected from YouTube (to be released) that spatialized by simulated room impulse responses (RIRs). Experiment results illustrate that our proposed multi-modal framework significantly outperforms single-modal and bi-modal speech separation approaches, while can still support real-time processing.
Keyword:
Lips
Feature extraction
Acoustics
Spectrogram
Visualization
Speech processing
Robustness
Target speech separation
speech enhancement
multi-modality fusion
deep learning
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Journal of Selected Topics in Signal Processing 封面图
IEEE Journal of Selected Topics in Signal Processing
IF:
13.7
论文数:
1.9K
被引数:
1.1W

机构

T
Tencent
学者数:
1.1K
论文数: 898
被引数: 5
P
peking university
学者数:
11.9W
论文数: 8.7W
被引数: 146