Return
Performance of State-of-the-Art Multimodal Large Language Models on an Image-Rich Radiology Board Examination: Comparison to Human Examinees
T
N
T
Y
H
M
S
Y
T
DOI:10.1016/j.acra.2025.10.051.png)
Abstract
En 中文
• Top-performing MLLMs (Gemini 2.5 Pro Preview: 76.0%; o3: 75.0%) exceeded the average human examinee score (72.9%) on the 2024 Japanese Radiology Board Examination (2nd stage). • Gemini 2.5 Pro Preview and Gemini 2.5 Flash Preview-thinking demonstrated statistically significant improvements in accuracy when provided with images (p = 0.035 and p = 0.019, respectively), effectively utilizing multimodal information. • Gemini models, despite their recent release, deliver top-class performance at a substantially lower cost (e.g., Gemini 2.5 Pro Preview at $1.25/$10.00 per 1 M input/output tokens vs. o3 at $10.00/$40.00), suggesting excellent cost-effectiveness for applications in radiology.
Keywords:
AI
Artificial Intelligence
API
Application Programming Interface
CT
Computed Tomography
JRS
Japan Radiological Society
LLM
Large Language Model
MLLM
Multimodal Large Language Model
MRI
Magnetic Resonance Imaging
MSK
Musculoskeletal
PACS
Picture Archiving and Communication Systems
PDF
Portable Document Format
RI
Radioisotope
RLHF
Reinforcement Learning from Human Feedback
SD
Standard Deviation
Large language models (LLMs)
Multimodal AI
Radiology Board Examination
Medical image analysis
Cost-effectiveness
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
3.9
Papers:
8.6K
Citations:
1.0W
