arrow
Return

A speech separation algorithm based on the comb-filter effect

delete2023-02-01
delete3
PRE
AI
张涛 cover
张涛 (Tao Zhang)
H
Heng Wang *
耿彦章 cover
耿彦章 (Yanzhang Geng)
X
Xin Zhao
L
Lingguo Kong
DOI:10.1016/j.apacoust.2022.109197delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In the field of speech separation, the traditional single-channel and multi-channel speech separation methods have made great progress. However, the accuracy of separation and automatic speech recogni-tion(ASR) rate are not yet satisfactory. With the development of neural networks, some scholars began to use deep learning to achieve speech separation. Although this kind of method improves the accuracy of speech separation, it also leads to the need for pre-training the model, higher computational complexity and reduced separation performance when the model does not match the mixed signal. This paper has conducted an in-depth study on the scene of multi-speaker separation, and proposed a new dual -channel speech separation algorithm based on the Comb-Filter Effect (CFE). The CFE is an effect that occurs when a signal passes through a first-order differential microphone(FDM) array. And this effect is discovered and exploited for the first time. By using this effect, this paper designed a new signal spec-trum estimation method that can realize accurate estimation of speech signal, and combined this method with traditional spectral subtraction to achieve the purpose of speech separation. Finally, this paper com-pared the proposed algorithm with the traditional FastICA-based algorithm and the fully-convolutional time-domain audio separation network(Conv-TasNet)-based algorithm. The results of simulation and comparison experiments show that the algorithm can effectively separate two-way speech signals while greatly reducing the computational complexity and has excellent robustness. In various situations, the proposed algorithm can obtain the Scale-Invariant Source-to-Noise Ratio improvement (SI-SNRi) of 9.19 dB on average. In addition, the Short-Time Objective Intelligibility (STOI) and Perceptual Evaluation of Speech Quality (PESQ) of the speech signal can be improved by an average of 33% and 70% or more respectively.(c) 2022 Elsevier Ltd. All rights reserved.
Keywords:
Differential microphone array
Comb -filter effect
Signal spectrum estimation
Speech separation

Journal

Applied Acoustics cover
Applied Acoustics
IF:
3.6
Papers:
7.3K
Citations:
1.7W

Organization

T
tianjin university
Scholars:
7.9W
Papers: 5.7W
Citations: 88