1
Return

Multi-Axial Analysis of Clinical Reasoning in Large Language Models: Inter-Verifier Disagreement and Its Implications for Automated Evaluation

delete2026-07-22
delete0
PRE
AI
H
Hyunjung Byun
D
D. T. Lee
M
Munyoung Jung
B
Beakcheol Jang *
DOI:10.1007/s10916-026-02440-ydelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Evaluating clinical reasoning in large language models (LLMs) poses two open challenges: reference-oriented semantic metrics do not directly assess whether a model’s stated diagnosis is supported by the evidence in its own justification, and the increasingly popular LLM-as-judge approach rests on a largely untested assumption—that independent verifier LLMs agree with one another. We assess three generator LLMs (HuatuoGPT-o1-8B, Meta-Llama-3.1-8B-Instruct, Meta-Llama-3.3-70B-Instruct) on 1,000 MIMIC-IV hospital-stay cases along four complementary axes (medical concept grounding, semantic similarity, semantic uncertainty, and evidence–conclusion coherence), with coherence judged independently by three frontier verifiers (Claude Sonnet 4.6, Gemini 2.5 Pro, GPT-5.4 mini). Two findings emerge. First, coherence reveals a dissociation that reference-oriented metrics do not capture: a model can score well on those axes yet still produce rationales that do not support its own conclusions. Second, inter-verifier agreement on coherence is consistently low (Fleiss’ $$\kappa$$ 0.087–0.223; disagreement 62.2%–74.3%), so the same rationale can be judged supported or unsupported depending on the verifier. A preliminary validation in which a physician adjudicated 50 cases echoed this: agreement with the physician varied across verifiers, underscoring that no single LLM reliably stands in for clinical assessment. Together, these results suggest a single LLM verifier lacks sufficient reliability to serve as a stand-alone judge of clinical reasoning at scale, and that structured human oversight remains essential. The unanimous-agreement tier offers a candidate for selective automation, but its clinical reliability remains to be confirmed in larger, multi-clinician adjudication studies.
Keywords:
Large language models
LLM-as-a-judge
Clinical reasoning
Hallucination
Medical AI evaluation
Evidence–conclusion coherence

Journal

Journal of Medical Systems cover
Journal of Medical Systems
IF:
5.7
Papers:
3.5K
Citations:
7.9K

Organization

C
College of Medicine
Scholars:
3.7K
Papers: 1.5K
Citations: 3
G
Graduate School of Information
Scholars:
19
Papers: 8
Citations: 0
G
Graduate School of Mechanical Engineering
Scholars:
8
Papers: 5
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers