arrow
Return

Automated speech emotion recognition model using feature-driven ensemble deep learning and puma optimization algorithm

delete2026-03-24
delete0
PRE
AI
A
Abeer S. Almogren
T
Taghreed Ali Alsudais
N
Nouf J. Aljohani
M
Mohammed Burhanur Rehman
F
Faten Derouez *
B
Basim Alamri
A
Ali M. Al-Sharafi
S
Somia A. Asklany
DOI:10.1016/j.eswa.2026.132213delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Speech is a prevalent and most common mode of human interaction. Recent consumer technologies such as home hubs and smartphones are increasingly equipped with advanced automatic speech recognition (ASR) systems to ease communication between humans and machines. ASR is converting speech signals into a sequence of words using a computer program. ASR is a pattern detection task in the computer science field, which is a challenging pattern recognition task in computer science. Traditionally, Gaussian Mixture Models (GMM) and Hidden Markov Models (HMM) have been the standard framework for speech recognition. However, with recent technological advances, deep learning (DL) and machine learning (ML) approaches have emerged as powerful tools for ASR. This paper presents an Automated Speech Recognition with Feature-Driven Ensemble Deep Learning Optimized by Puma Algorithm (ASR-FEDLOPA) model. The main goal of the ASR-FEDLOPA technique relies on enhancing the classification process for automated speech recognition using ensemble and feature models. The pre-processing stage is applied at dual levels, such as noise removal using anisotropic diffusion and zero padding. Furthermore, the ASR-FEDLOPA model employs three crucial features, tonal centroid, spectral contrast, and chroma, for effective feature extraction. For the classification process, the ensemble models, namely the Elman neural network (ENN), bidirectional recurrent neural network (BiRNN), and deep belief network (DBN), are employed. Finally, the hyperparameter selection of ensemble models is performed by implementing the Puma optimization operator (POO) model. The experimentation evaluation of the ASR-FEDLOPA approach is examined under the EMODB and RAVDESS datasets. The comparison analysis of the ASR-FEDLOPA approach portrayed superior accuracy values of 95.74% and 96.60% under dual datasets.
Keywords:
speech recognition
deep learning
ensemble models
feature extraction
puma optimization algorithm

Journal

Expert Systems with Applications cover
Expert Systems with Applications
IF:
7.5
Papers:
2.9W
Citations:
10.2W

Organization

U
University of Jeddah
Scholars:
2.3K
Papers: 2.6K
Citations: 3.6K
K
King Faisal University
Scholars:
4.4K
Papers: 4.4K
Citations: 5.0K
K
King Saud University
Scholars:
3.4W
Papers: 3.8W
Citations: 815
K
King Abdulaziz University
Scholars:
1.9W
Papers: 1.9W
Citations: 3.3W
N
Northern Border University
Scholars:
933
Papers: 739
Citations: 1.5K
P
princess nourah bint abdulrahman university
Scholars:
991
Papers: 1.1K
Citations: 0
U
University of Bisha
Scholars:
941
Papers: 1.0K
Citations: 1.2K
K
king khalid university
Scholars:
1.7K
Papers: 1.4K
Citations: 0
researcher View more organizations