arrow
Return

Improving Mandarin ASR Performance Through Multimodality

delete2025-11-18
delete0
delete
OA
AI
R
Rui Jiang
Y
Yang Zhao
伏晓 (Xiao Fu)
J
Jizhong Zhao *
DOI:10.3390/app152212224delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
In the context of Internet of Things (IoT) applications, accurate and efficient speech recognition is essential for enabling seamless voice-based interactions and control. Mandarin ASR, in particular, presents unique challenges due to the ideographic nature of the Chinese language, where recognition results are not directly correlated with pronunciation. Pinyin, as a representation of Chinese character pronunciation, has an intrinsic connection with Chinese characters, making it a valuable tool for enhancing ASR performance. This paper proposes a multimodal ASR neural network that combines pinyin data from the text modality and speech data from the audio modality as shared inputs to the ASR model. Specifically, the system processes the speech input through a preprocessed WeNet to generate pinyin text, which is then enhanced using a label denoising algorithm to improve its accuracy. The proposed text-acoustic multimodal ASR model improves the overall speech recognition performance by approximately 4%, making it more suitable for IoT applications that require high accuracy in voice commands and interactions.
Keywords:
automatic speech recognition
multimodal
label denoise
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

A
Applied Sciences-Basel
IF:
2.5
Papers:
7.3K
Citations:
4

Organization

X
Xi'an Jiaotong University
Scholars:
1.2W
Papers: 4.4K
Citations: 8.4W