arrow
Return

Multimodal prototypical network for interpretable sentiment classification

delete2025-10-26
delete0
delete
OA
AI
C
Chenguang Song *
柯超 (Chao Ke)
B
Bingjing Jia
Y
Yiqing Shen
DOI:10.1038/s41598-025-19850-6delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Recent advances in sentiment analysis have primarily focused on fusing multimodal information from video data, including visual, acoustic, and textual features, across temporal sequences. While great effort has been made to integrate or fuse information across modalities, less is known about the extent to which temporal segments contribute to model decisions. In addition, current interpretable methods, such as prototype networks, are primarily designed for uni-modal analysis and fail to handle the complex interactions between multiple modalities and temporal dependencies inherent in video data. To address the challenges, we propose MultiModal Prototypical Networks (MMPNet), which extends prototype-based interpretability to multimodal sentiment classification. Specifically, MMPNet can identify contributions of time-level features and leverage them to explain why a particular prediction was made, while also helping to find the relative importance of modality-level features. Experimental results show that MMPNet outperforms existing methods by 2.9% and 1.6% in accuracy on CMU-MOSI and CMU-MOSEI respectively, and achieves better interpretability.
Keywords:
Multimodal sentiment analysis
Multimodal prototypical networks
Interpretability

Journal

Scientific Reports cover
Scientific Reports
IF:
3.9
Papers:
27.4W
Citations:
83.5W

Organization

A
Anhui Science and Technology University
Scholars:
1.3K
Papers: 362
Citations: 1.5K
J
Johns Hopkins University
Scholars:
10.2W
Papers: 8.8W
Citations: 13.0W