Return
Automated speech emotion recognition model using feature-driven ensemble deep learning and puma optimization algorithm
DOI:10.1016/j.eswa.2026.132213.png)
Abstract
En 中文
Speech is a prevalent and most common mode of human interaction. Recent consumer technologies such as home hubs and smartphones are increasingly equipped with advanced automatic speech recognition (ASR) systems to ease communication between humans and machines. ASR is converting speech signals into a sequence of words using a computer program. ASR is a pattern detection task in the computer science field, which is a challenging pattern recognition task in computer science. Traditionally, Gaussian Mixture Models (GMM) and Hidden Markov Models (HMM) have been the standard framework for speech recognition. However, with recent technological advances, deep learning (DL) and machine learning (ML) approaches have emerged as powerful tools for ASR. This paper presents an Automated Speech Recognition with Feature-Driven Ensemble Deep Learning Optimized by Puma Algorithm (ASR-FEDLOPA) model. The main goal of the ASR-FEDLOPA technique relies on enhancing the classification process for automated speech recognition using ensemble and feature models. The pre-processing stage is applied at dual levels, such as noise removal using anisotropic diffusion and zero padding. Furthermore, the ASR-FEDLOPA model employs three crucial features, tonal centroid, spectral contrast, and chroma, for effective feature extraction. For the classification process, the ensemble models, namely the Elman neural network (ENN), bidirectional recurrent neural network (BiRNN), and deep belief network (DBN), are employed. Finally, the hyperparameter selection of ensemble models is performed by implementing the Puma optimization operator (POO) model. The experimentation evaluation of the ASR-FEDLOPA approach is examined under the EMODB and RAVDESS datasets. The comparison analysis of the ASR-FEDLOPA approach portrayed superior accuracy values of 95.74% and 96.60% under dual datasets.
Keywords:
speech recognition
deep learning
ensemble models
feature extraction
puma optimization algorithm
Journal
IF:
7.5
Papers:
2.9W
Citations:
10.2W

