arrow
返回

Multi-scale decomposition based supervised single channel deep speech enhancement

delete2020-10-01
delete16
PRE
AI
N
Nasir Saleem *
M
Muhammad Irfan Khattak
DOI:10.1016/j.asoc.2020.106666delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Speech signals reaching our ears are in general contaminated by the background noise distortion which is detrimental to both speech quality and intelligibility. In this paper, we propose a nonlinear multi-scale decomposition-based deep speech enhancement method to improve the quality and intelligibility of the contaminated speech. In the proposed method, we have applied Hurst exponent-based Empirical Mode Decomposition (HEMD) to the noisy speech and obtained a set of intrinsic mode functions (IMFs) and a residual. The Deep Neural Networks (DNNs) are trained for each of the extracted IMF and residual to learn a non-linear mapping with a deep hidden structure to construct a time-frequency mask. We have formulated three deep speech enhancement structures, established on three time-frequency masks comprised of Ideal Ratio Mask (IRM), Ideal Binary Mask (IBM), and Phase Sensitive Mask (PSM). Background noise also degrades the original phase of the clean speech; therefore, introduces perceptual disturbance which leads to negative impacts on the speech quality and intelligibility. To avoid speech quality and intelligibility degradations, an iterative procedure is adopted to compensate the phase during noisy backgrounds. Nonlinear Mel-scale weighted MSE (LMW-MSE) is used as a loss function during network training, and computed the gradients which are based on the perceptually motivated nonlinear frequency scale. Usually, the output features of the conventional deep neural networks are over-smoothed which deteriorates the quality of the speech. To alleviate over-smoothness; frequency-independent spectral variance equalization is applied as a post-filtering method. The performance of the proposed deep enhancement methods is extensively evaluated and compared to the DNNs established on same time-frequency mask in various adverse noisy environments. The results have demonstrated that the proposed deep speech enhancement performed better in terms of the perceived speech quality and intelligibility. (C) 2020 Elsevier B.V. All rights reserved.
Keyword:
DNN
EMD
Hurst exponent
Intelligibility
Speech enhancement
Objective loss function
Spectral variance equalization
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Applied Soft Computing 封面图
Applied Soft Computing
IF:
6.6
论文数:
1.4W
被引数:
4.8W

机构

U
University of Engineering and Technology Peshawar
学者数:
843
论文数: 712
被引数: 1.3K
引用论文

引用论文

Circulating gangliosides of breast‐cancer patients
err2006-07-17
err0
PREAI
errDouglas A. Wiesner; Charles C. Sweeley
err分享
err收藏
err分享
err收藏
Medication Use Among Inner-City Patients After Hospital Discharge: Patient-Reported Barriers and Solutions
err2008-05-01
err0
PREAI
errSunil Kripalani; Laura E. Henderson; Terry A. Jacobson; Viola Vaccarino
err分享
err收藏
Model-Based Speech Enhancement for Intelligibility Improvement in Binaural Hearing Aids
err2019-01-01
err24
errOAAI
errKavalekalam, Mathew Shaji; Nielsen, Jesper Kjaer; Boldt, Jesper Bunsow; Christensen, Mads Graesboll
err分享
err收藏
学者 查看更多内容