Return
Enhancing Vision-Language Model with Pretraining for Reasoning Medical Applications
Y
S
J
J
C
J
DOI:10.1016/j.compmedimag.2026.102765.png)
Abstract
En 中文
• We propose a Multi-modal Medical Reasoning Model (MMRM) that performs step-by-step diagnostic reasoning directly from medical images, enabling interpretable and traceable clinical decisions. • An Ortho-Enhanced Pretraining Framework improves the visual encoder’s ability to capture detailed semantic features, thus enhancing the quality of multi-modal representations. • Black-box Knowledge Distillation transfers Chain-of-Thought (CoT) reasoning strategies and domain-specific medical knowledge from a powerful teacher model to a smaller student language model. • A dedicated multi-modal CoT medical dataset integrates images, clinical questions, reasoning traces, and answers, allowing MMRM to generate reasoning explanations grounded in both visual and textual information. • This comprehensive three-stage training strategy enhances transparency and reliability in clinical decision support systems, contributing to the development of explainable AI in healthcare.
Keywords:
Multi-modal Medical Reasoning
Vision-Language Model
Chain-of-Thought Reasoning
Medical Image Analysis
Explainable AI
Journal
IF:
4.9
Papers:
2.4K
Citations:
5.0K

