arrow
Return

Controllable Conformer for Speech Enhancement and Recognition

delete2025-01-01
delete0
PRE
AI
Z
Zilu Guo
杜俊 (Jun Du) *
S
Sabato Marco Siniscalchi
J
Jia Pan
刘青锋 cover
刘青锋 (Qingfeng Liu)
DOI:10.1109/LSP.2024.3505794delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
We propose a novel approach to speech enhancement, termed Controllable ConforMer for Speech Enhancement (CCMSE), which leverages a Conformer-based architecture integrated with a control factor embedding module. Our method is designed to optimize speech quality for both human auditory perception and automatic speech recognition (ASR). It is observed that while mild denoising typically preserves speech naturalness, stronger denoising can improve human auditory tasks but often at the cost of ASR accuracy due to increased distortion. To address this, we introduce an algorithm that balances these trade-offs. By utilizing differential equations to interpolate between outputs at varying levels of denoising intensity, our method effectively combines the robustness of mild denoising with the clarity of stronger denoising, resulting in enhanced speech that is well-suited for both human and machine listeners. Experimental results on the CHiME-4 dataset validate the effectiveness of our approach.
Keywords:
Noise reduction
Speech enhancement
Signal to noise ratio
Estimation
Speech recognition
Noise measurement
Schedules
Differential equations
Decoding
Computer architecture
robust speech recognition
controllable denoising intensity
ASR

Journal

IEEE Signal Processing Magazine cover
IEEE Signal Processing Magazine
IF:
9.6
Papers:
1.1W
Citations:
1.7W

Organization

U
university of science & technology of china, cas
Scholars:
3.2W
Papers: 2.7W
Citations: 74
C
chinese academy of sciences
Scholars:
56.1W
Papers: 44.8W
Citations: 704