1
Return

MMA++: Effective Multi-Modal Adaptation for Vision-Language Models

delete2026-05-25
delete0
PRE
AI
L
Lingxiao Yang
张洳源 cover
张洳源 (Ru‐Yuan Zhang)
Y
Yanchen Wang
X
Xiaohua Xie
DOI:10.1109/tpami.2026.3691448delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Large scale pre-trained Vision-Language Models (VLMs) have shown good generalization capabilities across diverse downstream tasks. However, adapting such large-scale models to few-shot generalization scenarios remains challenging due to the trade-off between preserving general knowledge and incorporating task-specific information. In this paper, we propose <b>MMA++</b>, an advanced and effective <b>M</b>ulti-<b>M</b>odal <b>A</b>dapter framework for parameter-efficient VLM adaptation. Unlike prior works that independently inject adapters into each modality or uniformly across layers, MMA++ performs a dataset-level analysis to identify discriminative and generalizable features, and selectively applies adapters to the higher layers of both vision and text encoders. To bridge the modality gap, we further propose a shared feature projection space that enhances alignment between modalities. Beyond architecture design, we identify the fusion scale <inline-formula><tex-math notation="LaTeX">$\alpha$</tex-math></inline-formula>—which controls the strength of adapter integration—as a key factor in few-shot generalization. We empirically and theoretically demonstrate that <inline-formula><tex-math notation="LaTeX">$\alpha$</tex-math></inline-formula> should not be static, but adapted based on training data size. To reduce the effort of tuning this value across different datasets, we propose the <i><inline-formula><tex-math notation="LaTeX">$\alpha$</tex-math><alternatives><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML"><mml:mi>α</mml:mi></mml:math><inline-graphic xlink:href="xie-ieq3-3691448.gif" xmlns:xlink="http://www.w3.org/1999/xlink"/></alternatives></inline-formula>-consistency</i> framework, consisting of: (1) a consistency training strategy under varying fusion scales; and (2) an <inline-formula><tex-math notation="LaTeX">$\alpha$</tex-math></inline-formula>-decoupling strategy that uses a larger fusion scale during training and a smaller one at inference to account for sample size mismatch. We evaluate MMA++ on a wide range of few-shot generalization tasks, including base-to-novel generalization, cross-dataset transfer, and domain generalization. Our method consistently achieves leading performance.
Keywords:
Multi-modal adapter
vision-language models
few-shot adaptation
scale-consistency training
train-test decoupling

Journal

IEEE Transactions on Pattern Analysis and Machine Intelligence cover
IEEE Transactions on Pattern Analysis and Machine Intelligence
IF:
18.6
Papers:
831
Citations:
9.8W

Organization

C
columbia university
Scholars:
4.4K
Papers: 1.9K
Citations: 2
S
Sun Yat-Sen University
Scholars:
7.8K
Papers: 2.1K
Citations: 0
P
peking university
Scholars:
11.5W
Papers: 8.6W
Citations: 146
Cited Papers

Cited Papers

Citing Papers

Citing Papers