Return
GLIEAM: A ViT-GPT Dual-Stream Syntax-Guided Fusion Model for Multimodal Sentiment Analysis and Semantic Communication
F
DOI:10.1142/S1469026826500100.png)
Abstract
En 中文
To realize the cross-modal semantic comprehension, the authors suggest a GlobalLocal Interactive Emotion Analysis Model, referred to as GLIEAM, to overcome the problem of semantic inconsistency and noise interference encountered during multimodal comprehension. The model has a ViT of global visual semantics and a Generative Pre-Trained Transformer (GPT) of contextual linguistic representations. GLIEAM is built upon syntax-based semantic enhancement and multi-head cross-attention fusion, such that the language and the vision modalities can be aligned in an appropriate manner. GLIEAM performs better than baselines like BERT, TomBERT and SalienCyBERT; test results indicate that it is better than the baselines by up to 3.6% and F1, and in three datasets (Twitter-2015, Twitter-2017) and domains (UMLS, LegalPP-10k, WN18RR). The BR-GG-DeepSC fusion also enhances contextual semantic robustness and generalization under low-SNR. The ablation results show that the ViT-GPT fusion and the syntactic parsing/graphic denoising modules complement each other to achieve cross-modal alignment and representation stability. Overall, GLIEAM proposes a noise-resilient, interpretable, and generalizable approach to cross-modal semantic understanding and provides groundwork for advanced applications in semantic communication and multimodal reasoning.
Keywords:
Cross-modal semantic understanding
multimodal deep learning
ViT-GPT fusion
global-local interaction
semantic communication
Journal
I
IF:
1.3
Papers:
24
Citations:
0
