arrow
Return

Image Sensor-Supported Multimodal Attention Modeling for Educational Intelligence

delete2025-09-10
delete0
delete
OA
AI
Y
Yanlin Chen
Y
Yingqiu Yang
Z
Z.-J. Lan
X
Xinyuan Chen
H
Haoyuan Zhan
L
Lingxi Yu
Y
Yan Zhan *
DOI:10.3390/s25185640delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
To address the limitations of low fusion efficiency and insufficient personalization in multimodal perception for educational intelligence, a novel deep learning framework is proposed that integrates image sensor data with textual and contextual information through a cross-modal attention mechanism. The architecture employs a cross-modal alignment module to achieve fine-grained semantic correspondence between visual features captured by image sensors and associated textual elements, followed by a personalized feedback generator that incorporates learner background and task context embeddings to produce adaptive educational guidance. A cognitive weakness highlighter is introduced to enhance the discriminability of task-relevant features, enabling explicit localization and interpretation of conceptual gaps. Experiments show the proposed method outperforms conventional fusion and unimodal baselines with 92.37% accuracy, 91.28% recall, and 90.84% precision. Cross-task and noise-robustness tests confirm its stability, while ablation studies highlight the fusion module’s +4.2% accuracy gain and the attention mechanism’s +3.8% recall and +3.5% precision improvements. These results establish the proposed method as a transferable, high-performance solution for next-generation adaptive learning systems, offering precise, explainable, and context-aware feedback grounded in advanced multimodal perception modeling.

Journal

No journal information available

Organization

P
peking university
Scholars:
11.8W
Papers: 8.7W
Citations: 146