arrow
返回

Monaural Speech Dereverberation Using Deformable Convolutional Networks

delete2024-01-01
delete2
PRE
AI
V
Vinay Kothapally
J
John H. L. Hansen *
DOI:10.1109/TASLP.2024.3358720delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Reverberation and background noise can degrade speech quality and intelligibility when captured by a distant microphone. In recent years, researchers have developed several deep learning (DL)-based single-channel speech dereverberation systems that aim to minimize distortions introduced into speech captured in naturalistic environments. A majority of these DL-based systems enhance an unseen distorted speech signal by applying a predetermined set of weights to regions of the speech spectrogram, regardless of the degree of distortion within the respective regions. Such a system might not be an ideal solution for dereverberation task. To address this, we present a DL-based end-to-end single-channel speech dereverberation system that uses deformable convolution networks (DCN) that dynamically adjusts its receptive field based on the degree of distortions within an unseen speech signal. The proposed system includes the following components to simultaneously enhance the magnitude and phase responses of speech, which leads to improved perceptual quality: (i) a complex spectrum enhancement module that uses multi-frame filtering technique to implicitly correct the phase response, (ii) a magnitude enhancement module that suppresses dominant reflections and recovers the formant structure using deep filtering (DF) technique, and (iii) a speech activity detection (SAD) estimation module that predicts frame-wise speech activity to suppress residuals in non-speech regions. We assess the performance of the proposed system by employing objective speech quality metrics on both simulated and real speech recordings from the REVERB challenge corpus. The experimental results demonstrate the benefits of using DCNs and multi-frame filtering for speech dereverberation task. We compare the performance of our proposed system against other signal processing (SP) and DL-based systems and observe that it consistently outperforms other approaches across all speech quality metrics.
Keyword:
Speech enhancement
monaural dereverberation
deformable convolutional networks
minimum variance distortionless response
deep filtering

期刊

I
IEEE-ACM Transactions on Audio Speech and Language Processing
IF:
5.1
论文数:
2.6K
被引数:
1.1W

机构

U
university of texas system
学者数:
18.5W
论文数: 15.6W
被引数: 210
引用论文

引用论文

Impact of Nitrogen Fertilization on Phytophthora cinnamomi Root-related Damage in Juglans regia Saplings氮肥施用对Phytophthora cinnamomi引起的Juglans regia幼苗根系损伤的影响
err2019-12-01
err0
errOAAI
errJaviera Morales; Ximena Besoain; Italo F. Cuneo; Alejandra Larach; Laureano Alvarado; Alejandro Cáceres-Mella; Sebastian Saa
err分享
err收藏
Microscopy and Electrical Properties of Ge/Ge Interfaces Bonded by Surface-Activated Wafer Bonding Technology
err2011-01-20
err0
PREAI
errKentaroh Watanabe; Kensuke Wada; Hidehiro Kaneda; Kensuke Ide; Masahiro Kato; Takehiko Wada
err分享
err收藏
学者 查看更多内容