返回
Prototype-based sample-weighted distillation unified framework adapted to missing modality sentiment analysis
DOI:10.1016/j.neunet.2024.106397.png)
摘要
En 中文
Missing modality sentiment analysis is a prevalent and challenging issue in real life. Furthermore, the heterogeneity of multimodality often leads to an imbalance in optimization when attempting to optimize the same objective across all modalities in multimodal networks. Previous works have consistently overlooked the optimization imbalance of the network in cases when modalities are absent. This paper presents a PrototypeBased Sample -Weighted Distillation Unified Framework Adapted to Missing Modality Sentiment Analysis (PSWD). Specifically, it fuses features with a more efficient transformer -based cross -modal hierarchical cyclic fusion module. Subsequently, we propose two strategies, namely sample -weighted distillation and prototype regularization network, to address the issues of missing modality and optimization imbalance. The sampleweighted distillation strategy assigns higher weights to samples that are located closer to class boundaries. This facilitates the obtaining of complete knowledge by the student network from the teacher's network. The prototype regularization network calculates a balanced metric for each modality, which adaptively adjusts the gradient based on the prototype cross -entropy loss. Unlike conventional approaches, PSWD not only connects the sentiment analysis study in the missing modality to the full modality, but the proposed prototype regularization network is not reliant on the network structure and can be expanded to more multimodal studies. Massive experiments conducted on IEMOCAP and MSP-IMPROV show that our method achieves the best results compared to the latest baseline methods, which demonstrates its value for application in sentiment analysis.
Keyword:
Multimodal sentiment analysis
Missing modality
Optimization imbalance
Knowledge distillation
Prototype network
期刊
IF:
6.3
论文数:
8.2K
被引数:
3.0W
机构
引用论文
Cross-modal distillation with audio-text fusion for fine-grained emotion classification using BERT and Wav2vec 2.0
NEUROCOMPUTING
IF6.5
VisdaNet: Visual Distillation and Attention Network for Multimodal Sentiment Classification
SENSORS
IF3.5

