arrow
返回

Multi-Attention Bottleneck for Gated Convolutional Encoder-Decoder-Based Speech Enhancement

delete2023-01-01
delete6
delete
OA
AI
N
Nasir Saleem *
T
Teddy Surya Gunawan
S
Sami Bourouis
A
Aymen Trigui
DOI:10.1109/ACCESS.2023.3324210delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Convolutional encoder-decoder (CED) has emerged as a powerful architecture, particularly in speech enhancement (SE), which aims to improve the intelligibility and quality and intelligibility of noise-contaminated speech. This architecture leverages the strength of the convolutional neural networks (CNNs) in capturing high-level features. Usually, the CED architectures use the gated recurrent unit (GRU) or long-short-term memory (LSTM) as a bottleneck to capture temporal dependencies, enabling a SE model to effectively learn the dynamics and long-term temporal dependencies in the speech signal. However, Transformers neural networks with self-attention effectively capture long-term temporal dependencies. This study proposes a multi-attention bottleneck (MAB) comprised of a self-attention Transformer powered by a time-frequency attention (TFA) module followed by a channel attention module (CAM) to focus on the important features. The proposed bottleneck (MAB) is integrated into a CED architecture and named MAB-CED. The MAB-CED uses an encoder-decoder structure including a shared encoder and two decoders, where one decoder is dedicated to spectral masking and the other is used for spectral mapping. Convolutional Gated Linear Units (ConvGLU) and Deconvolutional Gated Linear Units (DeconvGLU) are used to construct the encoder-decoder framework. The outputs of two decoders are coupled by applying coherent averaging to synthesize the enhanced speech signal. The proposed speech enhancement is examined using two databases, VoiceBank+DEMAND and LibriSpeech. The results show that the proposed speech enhancement outperforms the benchmarks in terms of intelligibility and quality at various input SNRs. This indicates the performance of the proposed MAB-CED at improving the average PESQ by 0.55 (22.85%) with VoiceBank+DEMAND and by 0.58 (23.79%) with LibriSpeech. The average STOI is improved by 9.63% (VoiceBank+DEMAND) and 9.78% (LibriSpeech) over the noisy mixtures.
Keyword:
Speech enhancement
Convolutional neural networks
Time-frequency analysis
Decoding
Noise measurement
Logic gates
Transformers
Encoding
Multi-attention
time-frequency attention
channel attention
transformer
speech enhancement
gated convolutional encoder-decoder

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

G
Gomal University
学者数:
878
论文数: 788
被引数: 13
I
International Islamic University Malaysia
学者数:
2.6K
论文数: 2.0K
被引数: 1.9K
S
sohar university
学者数:
392
论文数: 560
被引数: 0
T
Taif University
学者数:
5.9K
论文数: 7.0K
被引数: 7.5K
K
King Khalid University
学者数:
1.1W
论文数: 1.3W
被引数: 1.5W
学者 查看更多机构
引用论文

引用论文

Circulating gangliosides of breast‐cancer patients
err2006-07-17
err0
PREAI
errDouglas A. Wiesner; Charles C. Sweeley
err分享
err收藏
err分享
err收藏
Individualizing Drug Dosage Regimens
err1993-10-01
err0
PREAI
errRoger W. Jelliffe; Alan Schumitzky; Michael Van Guilder; Min Liu; Lorinda Hu; Pascal Maire; Pilar Gomis; Xavier Barbaut; Babak Tahani
err分享
err收藏
Glance and gaze: A collaborative learning framework for single-channel speech enhancement
err2022-02-01
err99
errOAAI
errLi, Andong; Zheng, Chengshi; Zhang, Lu; Li, Xiaodong
err分享
err收藏
学者 查看更多内容