arrow
Return

Enhancing Vision-Language Model with Pretraining for Reasoning Medical Applications

delete2026-04-22
delete0
PRE
AI
Y
Yu Zhang
S
Shuihua Wang *
J
Jia Meng
J
Jiahe Hou
C
Christopher Overton
J
John Moraros
DOI:10.1016/j.compmedimag.2026.102765delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• We propose a Multi-modal Medical Reasoning Model (MMRM) that performs step-by-step diagnostic reasoning directly from medical images, enabling interpretable and traceable clinical decisions. • An Ortho-Enhanced Pretraining Framework improves the visual encoder’s ability to capture detailed semantic features, thus enhancing the quality of multi-modal representations. • Black-box Knowledge Distillation transfers Chain-of-Thought (CoT) reasoning strategies and domain-specific medical knowledge from a powerful teacher model to a smaller student language model. • A dedicated multi-modal CoT medical dataset integrates images, clinical questions, reasoning traces, and answers, allowing MMRM to generate reasoning explanations grounded in both visual and textual information. • This comprehensive three-stage training strategy enhances transparency and reliability in clinical decision support systems, contributing to the development of explainable AI in healthcare.
Keywords:
Multi-modal Medical Reasoning
Vision-Language Model
Chain-of-Thought Reasoning
Medical Image Analysis
Explainable AI

Journal

Computerized Medical Imaging and Graphics cover
Computerized Medical Imaging and Graphics
IF:
4.9
Papers:
2.4K
Citations:
5.0K

Organization

X
Xi'an Jiaotong-Liverpool University
Scholars:
443
Papers: 257
Citations: 5.4K
U
University of Liverpool
Scholars:
2.8W
Papers: 2.5W
Citations: 3.5W