arrow
返回

A Parallel-Data-Free Speech Enhancement Method Using Multi-Objective Learning Cycle-Consistent Generative Adversarial Network

delete2020-01-01
delete36
PRE
AI
Y
Yang Xiang
C
Changchun Bao *
DOI:10.1109/TASLP.2020.2997118delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Recently, deep neural networks (DNNs) have become the mainstream strategy for speech enhancement task because it can achieve the higher speech quality and intelligibility than the traditional methods. However, these DNN-based methods always need a large number of parallel corpus consisting of clean speech and noise to produce noisy data for the training of the DNN in order to improve the generalization of the network. As a result, this implies that many noisy speech signals that are collected in real environment cannot be used to train the DNN because of the lack of corresponding clean speech and noise. Additionally, as we know, noise varies with the time and scenario, so we cannot obtain parallel speech and noise due to infinite noise data and some limited speech data. Thus, the network training with unparallel speech and noise data is essential for the generalization of the network. To address this problem, we propose a novel parallel-data-free speech enhancement method, in which the cycle-consistent generative adversarial network (CycleGAN) and multi-objective learning are employed. Our method is also able to make best use of the benefits of multi-objective learning. On the training stage, we utilize two different encoders to encode the features of clean speech and noisy speech, respectively. Then, two forward generators are immediately used to predict the ideal time-frequency (T-F) mask and log-power spectrum (LPS) of clean speech. Two inverse generators are applied to map the magnitude spectrum (MS) and LPS of noisy speech, respectively. In addition, four discriminators are used to distinguish the real speech features from the generated features. Two encoders, four generators and four discriminators are simultaneously trained by using adversarial, identity-mapping, latent similarity and cycle-consistent loss. On the test stage, we directly utilize the forward generators and encoders to acquire the enhanced speech. The experimental results indicate that the proposed approach is able to achieve the better speech enhancement performance than the reference methods. Moreover, the proposed method is also effective to improve speech quality and intelligibility when the networks are trained under the parallel data.
Keyword:
Speech enhancement
Noise measurement
Generative adversarial networks
Generators
Task analysis
Gallium nitride
Cycle-consistent adversarial network
deep neural networks
multi-objective learning
non-parallel data
speech enhancement
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

I
IEEE-ACM Transactions on Audio Speech and Language Processing
IF:
5.1
论文数:
2.6K
被引数:
1.1W

机构

B
Beijing University of Technology
学者数:
2.8W
论文数: 2.1W
被引数: 2.7W
引用论文

引用论文

An Overview of Noise-Robust Automatic Speech Recognition
err2014-04-01
err421
PREAI
errLi, Jinyu; Deng, Li; Gong, Yifan; Haeb-Umbach, Reinhold
err分享
err收藏
Upregulation of GRP78 and GRP94 and Its Function in Chemotherapy Resistance to VP-16 in Human Lung Cancer Cell Line SK-MES-1
err2009-07-20
err0
PREAI
errLichuan Zhang; Siyan Wang; Wangtao; Yanyan Wang; Jiarui Wang; Li Jiang; Sheng Li; Xiujuan Hu; Qi Wang
err分享
err收藏
Microscopy and Electrical Properties of Ge/Ge Interfaces Bonded by Surface-Activated Wafer Bonding Technology
err2011-01-20
err0
PREAI
errKentaroh Watanabe; Kensuke Wada; Hidehiro Kaneda; Kensuke Ide; Masahiro Kato; Takehiko Wada
err分享
err收藏
err分享
err收藏
err分享
err收藏
学者 查看更多内容