arrow
Return

MULTI-CLASS AUTOMATED SPEECH LANGUAGE RECOGNITION USING NATURAL LANGUAGE PROCESSING WITH OPTIMAL DEEP LEARNING MODEL

delete2025-01-31
delete0
PRE
AI
R
Reema G. AL-anazi
H
Hamed Alqahtani
M
Muhammad Swaileh A. Alzaidi
M
Meshari Huwaytim Alanazi *
H
Hanan Al Sultan
A
Amal F. Alrowaily
J
Jawhara Aljabri
A
Assal A. M. Alqudah
DOI:10.1142/S0218348X25400213delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
With technological development, human-computer interaction (HCI) has improved, and spoken communication among machines and humans is one solution to enhance and expedite this process. Researchers have recently explored several systems to improve speech and speaker recognition performance in recent decades. A crucial threat in HCI is developing models that can effectually listen and respond like humans. It resulted in the development of the automated speech emotion recognition (SER) method, which can recognize various emotional classes by electing and extracting effectual features from speech signals. The fundamental problem of automated speech detection is the considerable variation in speech signals because of distinct speakers, language differences, speech differences, contents and acoustic conditions, voice modulation differences based on age and gender. With enhancements in deep learning (DL) and the affordability of computational resources, specifically graphical processing units (GPUs), research underwent a paradigm shift. Therefore, this study develops a multi-class automated speech language recognition using natural language processing with optimal deep learning (MASLR-NLPODL) technique. The MASLR-NLPODL technique intends to accomplish the efficient identification of different spoken languages. In the MASLR-NLPODL technique, the initial preprocessing technique involves windowing, frame blocking, and pre-emphasis block. Next, an adaptive time-frequency feature extractor approach utilizing the discrete fractional Fourier transform (DFrFT) was applied, which can be attained by extending the discrete Fourier transform (DFT) with eigenvectors. An improved Harris hawks optimization (IHHO) technique can be employed to select effectual features. Moreover, the classification of spoken languages can be performed by the gated recurrent unit (GRU) model. Finally, the salp swarm algorithm (SSA)-based hyperparameter selection process is involved in enhancing the performance of the GRU model. The design of the IHHO-based feature selection and SSA-based hyperparameter tuning process demonstrates the novelty of the work. The performance evaluation of the MASLR-NLPODL technique takes place under the VoxForge Dataset. The experimental validation of the MASLR-NLPODL technique exhibited a superior accuracy outcome of 96.40% over existing techniques.
Keywords:
Speech Language Recognition
Human-Computer Interaction
Hyperparameter Selection
Harris Hawks Optimization
NLP
Deep Learning

Journal

F
Fractals-Complex Geometry Patterns and Scaling in Nature and Society
IF:
2.9
Papers:
2.8K
Citations:
5.6K

Organization

K
King Faisal University
Scholars:
4.4K
Papers: 4.4K
Citations: 5.0K
K
King Saud University
Scholars:
3.4W
Papers: 3.8W
Citations: 815
P
Princess Nourah bint Abdulrahman University
Scholars:
7.9K
Papers: 9.3K
Citations: 10
K
King Saud bin Abdulaziz University for Health Sciences
Scholars:
5.4K
Papers: 2.8K
Citations: 2.3K
M
ministry of national guard - health affairs
Scholars:
1.3K
Papers: 788
Citations: 0
N
northern border university
Scholars:
1.8K
Papers: 2.1K
Citations: 2
K
King Khalid University
Scholars:
1.1W
Papers: 1.3W
Citations: 1.5W
U
University of Tabuk
Scholars:
4.3K
Papers: 4.1K
Citations: 3.5K
researcher View more organizations