arrow
Return

Multi-level visual-textual alignment transformer for multimodal aspect-based sentiment analysis

delete2025-10-25
delete0
PRE
AI
H
Hongxin Li
B
Bin Gao
L
Linlin Li
Y
Yutong Li
刘树田 (Shutian Liu)
Z
Zhengjun Liu
DOI:10.1016/j.eswa.2025.130148delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• We propose an MVAT approach for multimodal aspect-based sentiment analysis. • Align image and text by minimizing the Wasserstein distance of tokens and patches. • Designed a dynamic multimodal gate mechanism to integrate features from all directions. • We designed a GCN-based Differential Transformer to enhance focus on target aspect terms.

Journal

Expert Systems with Applications cover
Expert Systems with Applications
IF:
7.5
Papers:
2.9W
Citations:
10.2W

Organization

H
harbin institute of technology
Scholars:
8.0W
Papers: 6.6W
Citations: 66
H
Heilongjiang University
Scholars:
8.4K
Papers: 5.2K
Citations: 6.8K