1
Return

Audio-guided articulatory distillation for multilingual visual speech recognition with large language model decoding

delete2026-08-12
delete0
delete
OA
AI
B
Bogdan Mocanu
R
Ruxandra Țapu *
DOI:10.1007/s11042-026-21853-ydelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Visual Speech Recognition (VSR) remains challenging in multilingual and low-resource settings due to visual ambiguity, limited annotated data, and weak cross-lingual generalization. This paper introduces an audio-guided distillation framework for multilingual VSR that exploits synchronized audio-visual speech during training while preserving visual-only inference. The proposed architecture consists of an audio-visual teacher that learns articulation-aware representations from aligned acoustic and video streams, and a visual-only student trained to approximate the teacher through representation- and decoder-level distillation. To improve transcription under ambiguous visual evidence, continuous speech representations are projected into the embedding space of a pretrained large language model through a lightweight adaptation module, enabling language-conditioned decoding without full LLM retraining. We further introduce RoVSR-II, an extended Romanian in-the-wild VSR corpus comprising approximately 250 h of audiovisual speech, designed to support evaluation in an underrepresented language. Experiments on mTEDx demonstrate consistent improvements over existing multilingual VSR methods across Latin-script languages in terms of Word Error Rate (WER) and Character Error Rate (CER). Additional evaluation on RoVSR-II shows that the proposed model supports zero-shot transfer to Romanian and substantially reduces both WER and CER after parameter-efficient supervised adaptation. Ablation results further confirm the contribution of LLM decoding, teacher initialization, and audio-guided distillation to visual-only recognition performance. © 2017 Elsevier Inc. All rights reserved.
Keywords:
Visual speech recognition
Audio-guided distillation
Large language models
Low-resource languages
Teacher–student learning

Journal

Multimedia Tools and Applications cover
Multimedia Tools and Applications
IF:
3
Papers:
1.9W
Citations:
3.2W

Organization

F
faculty of etti
Scholars:
3
Papers: 1
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers