Return
Multi-level visual-textual alignment transformer for multimodal aspect-based sentiment analysis
DOI:10.1016/j.eswa.2025.130148.png)
Abstract
En 中文
• We propose an MVAT approach for multimodal aspect-based sentiment analysis. • Align image and text by minimizing the Wasserstein distance of tokens and patches. • Designed a dynamic multimodal gate mechanism to integrate features from all directions. • We designed a GCN-based Differential Transformer to enhance focus on target aspect terms.
Journal
IF:
7.5
Papers:
2.9W
Citations:
10.2W

