arrow
返回

Learning Deep Direct-Path Relative Transfer Function for Binaural Sound Source Localization

delete2021-01-01
delete16
delete
OA
AI
B
Bing Yang
刘宏 封面图
刘宏 (Hong Liu) *
X
Xiaofei Li
DOI:10.1109/TASLP.2021.3120641delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Direct-path relative transfer function (DP-RTF) refers to the ratio between the direct-path acoustic transfer functions of two microphone channels. Though DP-RTF fully encodes the sound spatial cues and serves as a reliable localization feature, it is often erroneously estimated in the presence of noise and reverberation. This paper proposes to learn DP-RTF with deep neural networks for robust binaural sound source localization. A DP-RTF learning network is designed to regress the binaural sensor signals to a real-valued representation of DP-RTF. It consists of a branched convolutional neural network module to separately extract the inter-channel magnitude and phase patterns, and a convolutional recurrent neural network module for joint feature learning. To better explore the speech spectra to aid the DP-RTF estimation, a monaural speech enhancement network is used to recover the direct-path spectrograms from the noisy ones. The enhanced spectrograms are stacked onto the noisy spectrograms to act as the input of the DP-RTF learning network. We train one unique DP-RTF learning network using many different binaural arrays to enable the generalization of DP-RTF learning across arrays. This way avoids time-consuming training data collection and network retraining for a new array, which is very useful in practical application. Experimental results on both simulated and real-world data show the effectiveness of the proposed method for direction of arrival (DOA) estimation in the noisy and reverberant environment, and a good generalization ability to unseen binaural arrays.
Keyword:
Location awareness
Feature extraction
Arrays
Speech enhancement
Spectrogram
Deep learning
Transfer functions
Direct-path relative transfer function
sound source localization
direction of arrival
deep neural network

期刊

I
IEEE-ACM Transactions on Audio Speech and Language Processing
IF:
5.1
论文数:
2.6K
被引数:
1.1W

机构

P
peking university
学者数:
11.9W
论文数: 8.7W
被引数: 146
W
westlake university
学者数:
5.3K
论文数: 3.7K
被引数: 8
引用论文

引用论文

err分享
err收藏
err分享
err收藏
The LOCATA Challenge: Acoustic Source Localization and Tracking
err2020-01-01
err98
errOAAI
errEvers, Christine; Loellmann, Heinrich W.; Mellmann, Heinrich; Schmidt, Alexander; Barfuss, Hendrik; Naylor, Patrick A.; Kellermann, Walter
err分享
err收藏
学者 查看更多内容