arrow
Return

Boosting Multi-Modal Large Language Model With Enhanced Visual Features

delete2025-12-16
delete0
PRE
AI
Y
Yiwei Ma
W
Weihuang Lin
Z
Zhibin Wang
J
Jiayi Ji
X
Xiaoshuai Sun
C
Chia-Wen Lin
R
Rongrong Ji
DOI:10.1109/TPAMI.2025.3644851delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Recent advancements in computer vision (CV) and large language models (LLMs) have spurred significant interest in multi-modal large language models (MLLMs), which aim to integrate visual and textual modalities for enhanced understanding and generation tasks. While much of the existing research focuses on optimizing projectors and LLMs to improve MLLM performance, a critical question remains underexplored: <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Has the full potential of visual features in MLLMs been realized?</i> To address this question, we identify two key limitations in current MLLM architectures and propose vMLLM, a vision-enhanced MLLM designed to fully leverage the capabilities of visual features. vMLLM introduces two novel components: the Multi-level Aggregation Module (MAM) and the Intra- and inter-modal Enhancement Module (IEM). The MAM aggregates multi-layer features from the vision encoder, capturing both high-level semantic information and low-level spatial details, thereby enriching the visual representation. The IEM enhances visual features through intra- and inter-modal interactions, effectively suppressing irrelevant information while amplifying task-relevant features, leading to more robust multimodal understanding. We conduct extensive experiments on multiple benchmarks, evaluating vMLLM across diverse settings, including different vision encoders, training dataset scales, and varying sizes of LLMs. Our results demonstrate that vMLLM consistently achieves significant performance improvements, validating its effectiveness in harnessing the potential of visual features. These findings highlight the importance of optimizing visual feature extraction and interaction mechanisms in MLLMs, paving the way for more advanced multimodal AI systems..
Keywords:
Multi-modal large language model
visual features
computer vision
large language model

Journal

IEEE Transactions on Pattern Analysis and Machine Intelligence cover
IEEE Transactions on Pattern Analysis and Machine Intelligence
IF:
18.6
Papers:
831
Citations:
9.8W

Organization

I
inf tech company
Scholars:
1
Papers: 1
Citations: 0
X
Xiamen University
Scholars:
5.3K
Papers: 1.7K
Citations: 6.2W
N
national tsing hua university
Scholars:
1.9K
Papers: 829
Citations: 0
researcher View more organizations