返回
Improving the loss function efficiency for speaker extraction using psychoacoustic effects
DOI:10.1016/j.apacoust.2021.108301.png)
摘要
En 中文
Speaker extraction aims to extract the speech signal of the speaker of interest from a mixture of two or more speakers. We propose a set of novel psychoacoustic-based loss functions that each can be used to optimize a stacked network of Bidirectional Long Short Term Memories (BLSTM) to imitate the human hearing system, which has an extraordinary ability to perceive and separate speech signals. To do this, we propose to use Mel and Gammatone filter banks as well as perceptual loudness and power-law of hearing effects in the loss functions of BLSTMs. The evaluation results on the Speech Separation Corpus (SSC) show that the proposed approach outperforms the baseline methods in terms of Short-Time Objective Intelligibility (STOI), Perceptual Evaluation of Speech Quality (PESQ) and Speech to Distortion Ratio (SDR). The proposed approach leads to an improvement of up to 0.276 in PESQ results compared to baseline methods. (C) 2021 Elsevier Ltd. All rights reserved.
Keyword:
Loss function
Auditory filter bank
Speaker extraction
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.6
论文数:
7.4K
被引数:
1.7W
机构
引用论文
Citizen science reveals the first occurrence of the greater white-toothed shrew Crocidura russula in Fennoscandia公民科学揭示了花白齿鼩Crocidura russula在芬兰-斯堪的纳维亚地区的首次出现。
Mammalia
IF0
没有更多内容

