1
Return

Multimodal Visual Speech Recognition for Under-Resource Languages via Cross-Modal Learning and Large Language Models

delete2026-01-01
delete0
PRE
AI
T
Tapu, Ruxandra *
M
Mocanu, Bogdan
C
Chiva, Ionut-Cosmin
DOI:10.59277/ROMJIST.2026.1.05delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This paper introduces a unified approach to multilingual visual speech recognition (VSR) that combines cross-modal phonetic modeling with large-scale language decoding to enable robust generalization across low-resource and previously unseen languages. The architecture within the approach includes a Cross-Modal Transcriber that encodes synchronized audio-visual speech inputs into a language-agnostic phoneme space via a fine-grained cross-attention mechanism. To bridge perception and language understanding, two decoding pathways are explored: (1) a modular configuration that maps phonetic sequences to text using a pretrained large language model (LLM), and (2) an end-to-end formulation in which fused visual features are projected into the LLM's embedding space via a lightweight adapter for direct transcription. Experimental evaluations on the mTEDx multilingual corpus show that the architecture surpasses state-of-the-art VSR models, achieving up to a 6% absolute improvement in WER across Latin-derived languages.
Keywords:
Cross-modal attention
large language models
multilingual learning
vi sual speech recognition

Journal

R
Romanian Journal of Information Science and Technology
IF:
3.9
Papers:
26
Citations:
442

Organization

I
imt - institut mines-telecom
Scholars:
7.4K
Papers: 6.3K
Citations: 5
I
institut polytechnique de paris
Scholars:
1.3W
Papers: 1.0W
Citations: 6
Cited Papers

Cited Papers

Citing Papers

Citing Papers