arrow
Return

Deep Learning System for Speech Command Recognition

delete2025-10-02
delete0
delete
OA
AI
D
Dejan Vujičić
Đ
Đorđe Damnjanović
D
Dušan Marković
Z
Zoran Stamenković *
DOI:10.3390/electronics14193793delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
We present a deep learning model for the recognition of speech commands in the English language. The dataset is based on the Google Speech Commands Dataset by Warden P., version 0.01, and it consists of ten distinct commands (“left”, “right”, “go”, “stop”, “up”, “down”, “on”, “off”, “yes”, and “no”) along with additional “silence” and “unknown” classes. The dataset is split in a speaker-independent manner, with 70% of speakers assigned to the training set and 15% to the test set and validation set. All audio clips are sampled at 16 kHz, with a total of 46 146 clips. Audio files are converted into Mel spectrogram representations, which are then used as input to a deep learning model composed of a four-layer convolutional neural network followed by two fully connected layers. The model employs Rectified Linear Unit (ReLU) activation, the Adam optimizer, and dropout regularization to improve generalization. The achieved testing accuracy is 96.05%. Micro- and macro-averaged precision, recall, and F1-score of 95% are reported to reflect class-wise performance, and a confusion matrix is also provided. The proposed model has been deployed on a Raspberry Pi 5 as a Fog computing device for real-time speech recognition applications.
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Electronics cover
Electronics
IF:
2.6
Papers:
9.6K
Citations:
4.7W

Organization

U
University of Kragujevac
Scholars:
3.0K
Papers: 2.0K
Citations: 1.7K
U
University of Potsdam
Scholars:
7.8K
Papers: 7.1K
Citations: 1.4W