arrow
返回

Phase Processing for Single-Channel Speech Enhancement

delete2015-03-01
delete200
PRE
AI
T
Timo Gerkmann *
M
Martin Krawczyk-Becker
J
Jonathan Le Roux
DOI:10.1109/MSP.2014.2369251delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
With the advancement of technology, both assisted listening devices and speech communication devices are becoming more portable and also more frequently used. As a consequence, users of devices such as hearing aids, cochlear implants, and mobile telephones, expect their devices to work robustly anywhere and at any time. This holds in particular for challenging noisy environments like a cafeteria, a restaurant, a subway, a factory, or in traffic. One way to making assisted listening devices robust to noise is to apply speech enhancement algorithms. To improve the corrupted speech, spatial diversity can be exploited by a constructive combination of microphone signals (so-called beamforming), and by exploiting the different spectro-temporal properties of speech and noise. Here, we focus on single-channel speech enhancement algorithms which rely on spectrotemporal properties. On the one hand, these algorithms can be employed when the miniaturization of devices only allows for using a single microphone. On the other hand, when multiple microphones are available, single-channel algorithms can be employed as a postprocessor at the output of a beamformer. To exploit the short-term stationary properties of natural sounds, many of these approaches process the signal in a time-frequency representation, most frequently the short-time discrete Fourier transform (STFT) domain. In this domain, the coefficients of the signal are complex-valued, and can therefore be represented by their absolute value (referred to in the literature both as STFT magnitude and STFT amplitude) and their phase. While the modeling and processing of the STFT magnitude has been the center of interest in the past three decades, phase has been largely ignored. In this article, we review the role of phase processing for speech enhancement in the context of assisted listening and speech communication devices. We explain why most of the research conducted in this field used to focus on estimating spectral magnitudes in the STFT domain, and why recently phase processing is attracting increasing interest in the speech enhancement community. Furthermore, we review both early and recent methods for phase processing in speech enhancement. We aim to show that phase processing is an exciting field of research with the potential to make assisted listening and speech communication devices more robust in acoustically challenging environments.
Keyword:
SPECTRAL MAGNITUDE ESTIMATION
TIME FOURIER-TRANSFORM
SIGNAL ESTIMATION
VOCODER
AUDIO
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Signal Processing Magazine 封面图
IEEE Signal Processing Magazine
IF:
9.6
论文数:
1.1W
被引数:
1.7W

机构

S
siemens ag
学者数:
5.6K
论文数: 4.6K
被引数: 3
引用论文

引用论文

Size dependency of hollow-cylinder stability
err1994-08-01
err0
PREAI
errHoek, van den; D.-J. Smit; A.P. Kooijman; de Ph.; C.J. Kenter; M. Khodaverdian
err分享
err收藏
err分享
err收藏
学者 查看更多内容