arrow
Return

A Survey on Efficient Vision-Language Models

delete2025-07-13
delete0
PRE
AI
G
Gaurav Shinde *
A
Anuradha Ravi
E
Emon Dey
S
Shadman Sakib
M
Milind Rampure
N
Nirmalya Roy
DOI:10.1002/widm.70036delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Vision-language models (VLMs) integrate visual and textual information, enabling a wide range of applications such as image captioning and visual question answering, making them crucial for modern AI systems. However, their high computational demands pose challenges for real-time applications. This has led to a growing focus on developing efficient vision-language models. In this survey, we review key techniques for optimizing VLMs on edge and resource-constrained devices. We also explore compact VLM architectures, frameworks, and provide detailed insights into the performance–memory trade-offs of efficient VLMs. Furthermore, we establish a GitHub repository at MPSC-GitHub to compile all surveyed papers, which we will actively update. Our objective is to foster deeper research in this area.
Keywords:
edge devices
efficient vision language models
multimodal models

Journal

Data Mining and Knowledge Discovery cover
Data Mining and Knowledge Discovery
IF:
4.3
Papers:
195
Citations:
6.0K

Organization

No organization information available