Return
Interpreting in-context learning in vision-language models for semantics–statistics disentanglement via out-of-distribution benchmark
Y
L
DOI:10.1016/j.neucom.2026.134676.png)
Abstract
En 中文
Understanding when and why in-context learning (ICL) helps vision-language models (VLMs) is central to building robust multimodal systems under distribution shift. We focus on out-of-distribution (OoD) settings because deployments rarely match training data, ICL is often the first line of adaptation, and OoD stress-tests whether models use semantics or spurious format cues. This study asks when ICL improves performance and how to separate task semantics from presentation statistics inside a VLM’s latent space. Prior work probes isolated concepts or treats ICL algorithmically (e.g., meta-optimization); however, it seldom provides a unified, VLM-specific account that disentangles semantics from statistics under OoD conditions. We propose Semantics–Statistics Disentanglement (S2D), formalizing projections onto semantic versus statistical subspaces, and derive two hypotheses: Semantic Boost (SB) and Textual Dominance Enhancement (TDE). We implement S2D with a paired, two-head probe with randomized controls and evaluate two LLaVA variants on one-shot OoD examples across six VQA datasets and a counterfactual-prompting setup on Hateful Memes. Results by hypothesis are as follows. For SB: accuracy gains concentrate where the zero-shot semantic signal is weak and become muted or slightly negative where it is already strong, and representation readouts show a higher semantic-over-statistical ratio after ICL. For TDE: counterfactual prompting raises F1 and lowers cross-label similarity on Hateful Memes, indicating reduced label bias and a shift toward text-appropriate dominance. In addition, a simple selector guided by the zero-shot state outperforms fixed strategies when zero-shot and ICL are comparable. Together, these findings offer a principled account of when ICL helps under distribution shift and yield actionable diagnostics for choosing context strategies that maximize semantic gain while limiting statistical amplification.
Keywords:
Vision-language model
In-context learning
Interpretability
Out-of-distribution generalization
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W
