1
Return

Evaluating Clinical Reasoning and Diagnostic Performance of Multimodal Large Language Models in Benign Hematology—A Comparative Analysis

delete2026-07-11
delete0
delete
OA
AI
Z
Zahra Rashid Khan
K
Kamran Khan
M
Maham Nayab
M
Muhammad Usman Akram
S
Sundas Ali
H
Huma Abdul Shakoor
A
Ayesha Haque
M
Meeral Ahmed *
DOI:10.1002/ima.70409delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Latest multimodal LLMs, capable of analyzing textual data and images. have potential application in decision support systems in diagnostics. Their performance and clinical reasoning abilities in benign hematology remain underexplored. This study evaluates and compares diagnostic performance and clinical reasoning of 3 LLM in classification of anemia, based on CBC triggers and peripheral blood smears, benchmarked with expert validated cases. A total of 270 LLM generated responses were evaluated for accuracy and clinical reasoning, using real world, expert validated anemia cases. LLMs (Gemini 2.5 Pro, Chat-GPT4o and Claude Opus 4 Pro) were presented with focused prompt, digital image and CBC parameters as multimodal input. LLM accuracy scores and rubric-derived clinical reasoning was calculated. Kruskal–Wallis H test was used to detect significant differences in clinical reasoning across LLMs, followed up with Dunn's Post Hoc Pairwise comparisons where indicated. Correlation between accuracy and clinical reasoning scores was evaluated using Spearman's correlation coefficient. Gemini 2.5 Pro showed the highest accuracy values (54%) among the three LLMs in Overall Accuracy, while Chat GPT 4o scored the lowest (38%). Class-wise accuracies showed Claude Opus Pro 4 (76%) as the most accurate for Iron Deficiency Anemia and Gemini 2.5 (50%) for Hemoglobinopathy. There was a significant difference among LLMs in diagnosis of Hemoglobinopathy (p = 0.016). A significant difference in clinical reasoning scores was noted across LLMs for Iron Deficiency Anemia (H statistic = 6.216, p = 0.0446). Post hoc pairwise comparisons revealed Gemini versus Claude pair to be statistically significant (p = 0.040). A similar observation was made across LLM scores for the Hemoglobinopathy class (H statistic = 6.743, p = 0.0343). Post hoc pairwise comparisons revealed Gemini versus GPT-4o pair to be statistically significant (p = 0.0390) Additionally, Spearman correlation coefficient for accuracy versus clinical reasoning scores demonstrated consistently strong correlations across the three anemia subtypes for Gemini 2.5 Pro and Claude Opus 4. (ρ scores between 0.79 and 0.83), with GPT 4o showing modest correlation. While Gemini 2.5 Pro and Claude Opus 4 Pro perform relatively better among the three LLMs, their accuracy and clinical reasoning scores show only partial readiness for integration in Hematology training/teaching and workflows. These values should, however, be viewed as indicative rather than conclusive due to limited dataset. Although the results do provide insight but should be interpreted as preliminary and hypothesis-generating.
Keywords:
anemia
artificial intelligence
clinical reasoning
diagnosis
hematology
large language model
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

International Journal of Imaging Systems and Technology cover
International Journal of Imaging Systems and Technology
IF:
2.5
Papers:
2.1K
Citations:
2.3K

Organization

Pakistan Institute of Medical Sciences cover
Pakistan Institute of Medical Sciences
Scholars:
250
Papers: 125
Citations: 136
A
atomic energy cancer hospital
Scholars:
2
Papers: 1
Citations: 0
L
loyola university chicago
Scholars:
926
Papers: 407
Citations: 0
N
national university of sciences and technology
Scholars:
679
Papers: 361
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers