返回
Multi-layer encoder-decoder time-domain single channel speech separation
DOI:10.1016/j.patrec.2024.03.020.png)
摘要
En 中文
With the emergence of more advanced separation networks, significant progress has been made in timedomain speech separation methods. These methods typically use a temporal encoder-decoder structure to encode speech feature sequences, thereby accomplishing the separation task. However, due to the limitation of traditional encoder-decoder structure, the separation performance decreases sharply when the encoded sequence is short, and when encoded sequence is sufficiently long, the separation performance improves, but which leads to an increase in computational complexity and training cost. Therefore, this paper compresses and reconstructs the speech feature sequence through a multi-layer convolution structure, and proposes a multilayer encoder-decoder time-domain speech separation model (MLED). In this model, our encoder-decoder structure can compress speech sequence to a short length while ensuring the separation performance does not decrease. And combined with our multi-scale temporal attention (MSTA) separation network, MLED achieves efficient and precise separation of short encoded sequences. Therefore, compared to previous advanced timedomain separation methods, our experiments show that MLED achieves competitive separation performance with smaller model size, lower computational complexity, and training cost.
Keyword:
Time-domain speech separation
Attention mechanism
Multi-layer encoder-decoder
Training cost
期刊
IF:
3.3
论文数:
7.9K
被引数:
1.6W
机构
引用论文
Solution structure of the Lewis x oligosaccharide determined by NMR spectroscopy and molecular dynamics simulations
Biochemistry
IF0
Neonatal caffeine administration causes a permanent increase in the dendritic length of prefrontal cortical neurons of rats
Synapse
IF0
Shape analysis of the neostriatum in frontotemporal lobar degeneration, Alzheimer's disease, and controls
NeuroImage
IF0
U2-Net: Going deeper with nested U-structure for salient object detectionU2-Net: 基于嵌套U结构的显著性目标检测
PATTERN RECOGNITION
IF7.6

