arrow
返回

Benchmarking large and small MLLMs

delete2025-10-30
delete1
delete
OA
AI
F
Feng, Xuelu *
Y
Y. Li
D
Dongdong Chen
M
Mei Gao
L
Liu, Mengchen
J
Junsong Yuan
C
Chunming Qiao
DOI:10.1007/s00138-025-01762-0delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Large multimodal language models (MLLMs) such as GPT-4V and GPT-4o have achieved remarkable advancements in understanding and generating multimodal content, showcasing superior quality and capabilities across diverse tasks. However, their deployment faces significant challenges, including slow inference, high computational cost, and impracticality for on-device applications. In contrast, the emergence of small MLLMs, exemplified by the LLava-series models and Phi-3-Vision, offers promising alternatives with faster inference, reduced deployment costs, and the ability to handle domain-specific scenarios. Despite their growing presence, the capability boundaries between large and small MLLMs remain underexplored. In this work, we conduct a systematic and comprehensive evaluation to benchmark both small and large MLLMs, spanning general capabilities such as object recognition, temporal reasoning, and multimodal comprehension, as well as real-world applications in domains like industry and automotive. Our evaluation reveals that small MLLMs can achieve comparable performance to large models in specific scenarios but lag significantly in complex tasks requiring deeper reasoning or nuanced understanding. Furthermore, we identify common failure cases in both small and large MLLMs, highlighting domains where even state-of-the-art models struggle. We hope our findings will guide the research community in pushing the quality boundaries of MLLMs, advancing their usability and effectiveness across diverse applications.
Keyword:
Multi-modal large language model
Evaluation
General capabilities
Application
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

M
Machine Vision and Applications
IF:
2.3
论文数:
95
被引数:
2.7K

机构

S
state university of new york (suny) system
学者数:
6.5W
论文数: 5.8W
被引数: 65
U
university at buffalo, suny
学者数:
1.2W
论文数: 9.5K
被引数: 9
引用论文

引用论文

SEED-Bench: Benchmarking Multimodal Large Language Models
err2024-06-16
err0
PREAI
errBohao Li; Yuying Ge; Yixiao Ge; Guangzhi Wang; Rui Wang; Ruimao Zhang; Ying Shan
err分享
err收藏
The Cityscapes Dataset for Semantic Urban Scene Understanding用于语义城市场景理解的Cityscapes数据集
err2016-06-01
err0
errOAAI
errMarius Cordts; Mohamed Omran; Sebastian Ramos; Timo Rehfeld; Markus Enzweiler; Rodrigo Benenson; Uwe Franke; Stefan Roth; Bernt Schiele
err分享
err收藏
Large Scale Visual Food Recognition
err2023-08-01
err0
errOAAI
errWeiqing Min; Zhiling Wang; Yuxin Liu; Mengjiang Luo; Liping Kang; Xiaoming Wei; Xiaolin Wei; Shuqiang Jiang
err分享
err收藏
Evaluating Count Prioritization Procedures for Improving Inventory Accuracy in Retail Stores
err2023-01-01
err0
errOAAI
errNicole DeHoratius; Andreas Holzapfel; Heinrich Kuhn; Adam J. Mersereau; Michael Sternbeck
err分享
err收藏
Scene Text Visual Question Answering
err2019-10-01
err0
errOAAI
errAli Furkan Biten; Ruben Tito; Andres Mafla; Lluis Gomez; Marcal Rusinol; C.V. Jawahar; Ernest Valveny; Dimosthenis Karatzas
err分享
err收藏
err分享
err收藏
Learning Deep Features for Discriminative Localization
err2016-06-01
err0
errOAAI
errBolei Zhou; Aditya Khosla; Agata Lapedriza; Aude Oliva; Antonio Torralba
err分享
err收藏
Are Large Vision Language Models up to the Challenge of Chart Comprehension and Reasoning
err2024-01-01
err0
PREAI
errIslam,Mohammed Saidul; Rahman,Raian; Masry,Ahmed; Laskar,Md Tahmid Rahman; Nayeem,Mir Tafseer; Hoque,Enamul
err分享
err收藏
学者 查看更多内容