arrow
Return

Visual speech recognition using compact hypercomplex neural networks

delete2024-10-01
delete1
PRE
AI
I
Iason-Ioannis Panagos *
G
Giorgos Sfikas
C
Christophoros Nikou
DOI:10.1016/j.patrec.2024.09.002delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Recent progress in visual speech recognition systems due to advances in deep learning and large-scale public datasets has led to impressive performance compared to human professionals. The potential applications of these systems in real-life scenarios are numerous and can greatly benefit the lives of many individuals. However, most of these systems are not designed with practicality in mind, requiring large-size models and powerful hardware, factors which limit their applicability in resource-constrained environments and other real-world tasks. In addition, few works focus on developing lightweight systems that can be deployed in such conditions. Considering these issues, we propose compact networks that take advantage of hypercomplex layers that utilize a sum of Kronecker products to reduce overall parameter demands and model sizes. We train and evaluate our proposed models on the largest public dataset for single word speech recognition for English. Our experiments show that high compression rates are achievable with a minimal accuracy drop, indicating the method's potential for practical applications in lower-resource environments. Code and models are available at https://github.com/jpanagos/vsr_phm.
Keywords:
Visual speech recognition
Lipreading
Hypercomplex multiplication

Journal

Pattern Recognition Letters cover
Pattern Recognition Letters
IF:
3.3
Papers:
7.9K
Citations:
1.6W

Organization

U
University of Ioannina
Scholars:
8.0K
Papers: 7.1K
Citations: 8.0K
U
University of West Attica
Scholars:
2.3K
Papers: 1.8K
Citations: 1.2K
Cited Papers

Cited Papers

Audio-Visual Speech and Gesture Recognition by Sensors of Mobile Devices
errSENSORS
IF3.5
err2023-02-17
err55
errOAAI
errRyumin, Dmitry; Ivanko, Denis; Ryumina, Elena
errShare
errSave
Pattern dynamics in a diffusive predator-prey model with hunting cooperations
err2020-01-01
err0
PREAI
errShuixian Yan; Dongxue Jia; Tonghua Zhang; Sanling Yuan
errShare
errSave
A review of recent advances in visual speech decoding
err2014-09-01
err133
PREAI
errZhou, Ziheng; Zhao, Guoying; Hong, Xiaopeng; Pietikainen, Matti
errShare
errSave
no more