Return
Multimodal Aspect-Based Sentiment Analysis With Plugin-Enhanced Large Language Models
DOI:10.1109/TNNLS.2025.3622470.png)
Abstract
En 中文
Multimodal aspect-based sentiment analysis (MABSA) is a challenging task that predicts sentiment polarity for specific aspect terms based on inputs across modalities. Existing approaches typically employ advanced visual and textual encoders to extract multimodal features and align them for MABSA prediction, yet they still face challenges in handling complex connections between multiple modalities. Recent bloom of large language models (LLMs), as well as their multimodal counterparts, has shown significant promise in various tasks, which offer a promising solution for MABSA, with potential limitations such as semantic mismatch between images and texts, and their high computational cost of fine-tuning for specific tasks. To address these limitations, in this article, we propose a novel plugin-based approach for MABSA, which uses plugins to encode key knowledge instances, such as salient objects in images and word relationships in texts, with an attentive graph convolutional network (A-GCN). We further utilize a memory-based hub to integrate the encoded multimodal knowledge and align the knowledge representations with the LLM, guiding it to better understand the intricate connections between modalities. We evaluate our approach on two benchmark MABSA datasets, which outperforms baselines and achieves state-of-the-art performance over existing studies. Further analysis shows that our approach enables efficient and scalable adaptation of multimodal LLMs to specific tasks, making it a promising solution for related tasks. The code is available at https://github.com/synlp/MABSA-LLMPlug
Keywords:
Knowledge
multimodal aspect-based sentiment analysis (MABSA)
multimodal large language models (LLMs)
plugins
Journal
IF:
8.9
Papers:
7.5K
Citations:
7.2W

