Return
Evaluating Clinical Reasoning and Diagnostic Performance of Multimodal Large Language Models in Benign Hematology—A Comparative Analysis
Z
K
M
M
S
H
A
M
DOI:10.1002/ima.70409.png)
Abstract
En 中文
Latest multimodal LLMs, capable of analyzing textual data and images. have potential application in decision support systems in diagnostics. Their performance and clinical reasoning abilities in benign hematology remain underexplored. This study evaluates and compares diagnostic performance and clinical reasoning of 3 LLM in classification of anemia, based on CBC triggers and peripheral blood smears, benchmarked with expert validated cases. A total of 270 LLM generated responses were evaluated for accuracy and clinical reasoning, using real world, expert validated anemia cases. LLMs (Gemini 2.5 Pro, Chat-GPT4o and Claude Opus 4 Pro) were presented with focused prompt, digital image and CBC parameters as multimodal input. LLM accuracy scores and rubric-derived clinical reasoning was calculated. Kruskal–Wallis H test was used to detect significant differences in clinical reasoning across LLMs, followed up with Dunn's Post Hoc Pairwise comparisons where indicated. Correlation between accuracy and clinical reasoning scores was evaluated using Spearman's correlation coefficient. Gemini 2.5 Pro showed the highest accuracy values (54%) among the three LLMs in Overall Accuracy, while Chat GPT 4o scored the lowest (38%). Class-wise accuracies showed Claude Opus Pro 4 (76%) as the most accurate for Iron Deficiency Anemia and Gemini 2.5 (50%) for Hemoglobinopathy. There was a significant difference among LLMs in diagnosis of Hemoglobinopathy (p = 0.016). A significant difference in clinical reasoning scores was noted across LLMs for Iron Deficiency Anemia (H statistic = 6.216, p = 0.0446). Post hoc pairwise comparisons revealed Gemini versus Claude pair to be statistically significant (p = 0.040). A similar observation was made across LLM scores for the Hemoglobinopathy class (H statistic = 6.743, p = 0.0343). Post hoc pairwise comparisons revealed Gemini versus GPT-4o pair to be statistically significant (p = 0.0390) Additionally, Spearman correlation coefficient for accuracy versus clinical reasoning scores demonstrated consistently strong correlations across the three anemia subtypes for Gemini 2.5 Pro and Claude Opus 4. (ρ scores between 0.79 and 0.83), with GPT 4o showing modest correlation. While Gemini 2.5 Pro and Claude Opus 4 Pro perform relatively better among the three LLMs, their accuracy and clinical reasoning scores show only partial readiness for integration in Hematology training/teaching and workflows. These values should, however, be viewed as indicative rather than conclusive due to limited dataset. Although the results do provide insight but should be interpreted as preliminary and hypothesis-generating.
Keywords:
anemia
artificial intelligence
clinical reasoning
diagnosis
hematology
large language model
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
2.5
Papers:
2.1K
Citations:
2.3K

