Return
MMA++: Effective Multi-Modal Adaptation for Vision-Language Models
L
Y
X
DOI:10.1109/tpami.2026.3691448.png)
Abstract
En 中文
Large scale pre-trained Vision-Language Models (VLMs) have shown good generalization capabilities across diverse downstream tasks. However, adapting such large-scale models to few-shot generalization scenarios remains challenging due to the trade-off between preserving general knowledge and incorporating task-specific information. In this paper, we propose <b>MMA++</b>, an advanced and effective <b>M</b>ulti-<b>M</b>odal <b>A</b>dapter framework for parameter-efficient VLM adaptation. Unlike prior works that independently inject adapters into each modality or uniformly across layers, MMA++ performs a dataset-level analysis to identify discriminative and generalizable features, and selectively applies adapters to the higher layers of both vision and text encoders. To bridge the modality gap, we further propose a shared feature projection space that enhances alignment between modalities. Beyond architecture design, we identify the fusion scale <inline-formula><tex-math notation="LaTeX">$\alpha$</tex-math></inline-formula>—which controls the strength of adapter integration—as a key factor in few-shot generalization. We empirically and theoretically demonstrate that <inline-formula><tex-math notation="LaTeX">$\alpha$</tex-math></inline-formula> should not be static, but adapted based on training data size. To reduce the effort of tuning this value across different datasets, we propose the <i><inline-formula><tex-math notation="LaTeX">$\alpha$</tex-math><alternatives><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mi>α</mml:mi></mml:math><inline-graphic xlink:href="xie-ieq3-3691448.gif" xmlns:xlink="http://www.w3.org/1999/xlink"/></alternatives></inline-formula>-consistency</i> framework, consisting of: (1) a consistency training strategy under varying fusion scales; and (2) an <inline-formula><tex-math notation="LaTeX">$\alpha$</tex-math></inline-formula>-decoupling strategy that uses a larger fusion scale during training and a smaller one at inference to account for sample size mismatch. We evaluate MMA++ on a wide range of few-shot generalization tasks, including base-to-novel generalization, cross-dataset transfer, and domain generalization. Our method consistently achieves leading performance.
Keywords:
Multi-modal adapter
vision-language models
few-shot adaptation
scale-consistency training
train-test decoupling
Journal
IF:
18.6
Papers:
831
Citations:
9.8W
