Return
Recognition of speech using deep learning-based optimization technique
DOI:10.1142/S0219691326500128.png)
Abstract
En 中文
Speech recognition from noisy speech signals remains a significant challenge in the field of human-computer interaction. Speech signals are often degraded by reverberation and background noise, leading to reduced recognition accuracy. Conventional methods typically rely on large volumes of training data and fail to generalize effectively under noisy conditions. Moreover, these approaches often exhibit limited robustness and reduced performance in adverse environments. In this work, a new technique called Fractional Super Bilterling Optimization with Random Multimodal Deep Learning-Convolutional Neural Network (FrSBO_RMDL-CNN) is proposed for speech recognition from unclear speech signals. At first, the unclear speech signal collected from the database is denoised by employing WaveNet denoising. Next, the voice enhancement is performed utilizing the Speech Enhancement and Language Model (SELM), which is trained using the Super Bilterling Optimization (SBiO). Here, the SBiO is developed by the integration of Bilterling Fish Optimization (BFO) and Superb Fairy-wren Optimization Algorithm (SFOA). After that, the speech word is segmented employing the Attentional Encoder-Decoder. Last, the speech recognition is performed using the RMDL-CNN, which is tuned by Fractional Super Bilterling Optimization (FrSBO). Furthermore, FrSBO_RMDL-CNN computed the maximum Negative Predictive Value (NPV), recognition accuracy and Positive Predictive Value (PPV) of 97.146%, 97.798% and 96.888%, respectively.
Keywords:
Speech
voice disorder
voice signal enhancement
fractional super bilterling optimization
deep learning
Journal
I
IF:
0.8
Papers:
39
Citations:
717

