1
Return

Beyond Accuracy: A Mixed-Methods Audit of Chain-of-Thought Failures in LLM-Based COVID-19 Vaccine Stance Detection

delete2026-08-08
delete0
delete
OA
AI
A
Andreas Praschk
V
Valentin Fischill-Neudeck
T
Thomas Caspari
H
Hans‐Peter Wiesinger *
DOI:10.1007/s10916-026-02449-3delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This mixed-methods study assessed whether reasoning-enabled large language models (LLMs) can classify stances towards COVID-19 vaccination on X (formerly Twitter) and whether model-generated chain-of-thought (CoT) summaries contain reasoning failures relevant to transparent and auditable public health applications. Zero-shot stance classification by o4-mini and Gemini 2.5 Flash (Gemini) was evaluated on 3,060 rehydrated COVID-19 vaccination tweets against human-annotated labels (positive, negative, neutral). We reported accuracy and macro-F1, measured CoT availability, and qualitatively analysed dual-error cases (tweets misclassified by both models) using Mayring’s content analysis guided by the FUTURE-AI framework. At each model’s best-performing setting, both models reached macro-F1 around 0.8, with o4-mini outperforming Gemini (accuracy 0.819 vs. 0.799, McNemar p = 0.0015; Δmacro-F1 = 0.020, 95% CI 0.008–0.032). Under the reasoning-intensive settings, CoT availability differed: Gemini returned a reasoning summary for all tweets, whereas o4-mini did so for 64.7%. Among 1,981 tweets with CoTs from both models, 295 (14.9%) were dual-errors; in 88.8%, both models produced the same wrong label, suggesting shared failure modes. Qualitatively, both models showed the same errors: target confusion (policy vs. vaccine), literal readings of sarcasm, and label–rationale mismatches, recurring across models despite their markedly different CoT lengths. Reasoning LLMs can therefore classify stance accurately, but their readiness for transparent public health applications depends on whether a CoT is available at all and whether it is coherent with the label it accompanies (label–rationale coherence). CoT availability, label–rationale coherence, and safeguards against systematic reasoning failures offer candidate explainability-readiness metrics, alongside accuracy, for trustworthy digital epidemiology.
Keywords:
Large language models
Chain-of-thought
Stance detection
Explainability
Vaccine hesitancy
Mixed-methods study

Journal

Journal of Medical Systems cover
Journal of Medical Systems
IF:
5.7
Papers:
3.5K
Citations:
7.9K

Organization

C
Center for Public Health and Healthcare Research
Scholars:
5
Papers: 3
Citations: 0
F
faculty of medicine
Scholars:
6.1K
Papers: 2.1K
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers