1
Return

Performance and limitations of four large language models in genetic counseling for thalassemia

delete2026-08-11
delete0
delete
OA
AI
W
WZ Wenfu Zhong †
J
JH Jingwen Huang †
M
Mengsi Wei
Q
Qingpeng Liang
Y
YY Yuanwu Yang
J
JC Jinrong Chen
J
Jinjiang Mao
秦颖 (Ying Qin)
Y
YS Yifan Sun
Y
YL Yishan Liang
DOI:10.3389/fdgth.2026.1902480delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In regions with high thalassemia prevalence; such as southern China and Southeast Asia; chronic shortages of professional genetic counseling resources have driven interest in large language models (LLMs) as auxiliary tools; yet their performance and safety boundaries in this setting remain uncharacterized. This single-center retrospective study evaluated four LLMs (ChatGPT-5.2 Thinking; DeepSeek-V3.2 chat; Gemini 3 Flash; and Grok 4.1 Fast) using 1; 080 standardized knowledge questions administered across five independent sessions and 150 real-world clinical cases scored by six senior experts across eight dimensions. All models exceeded 90% accuracy on single-choice and true-false questions. Between-model differences were most pronounced in multiple-choice questions; where ChatGPT-5.2 Thinking achieved the highest accuracy (87.28% ± 1.69%); significantly outperforming Grok 4.1 Fast (72.06% ± 1.69%; P < 0.001). DeepSeek-V3.2 chat showed lower cross-session consistency than the other three models. In clinical case analysis; all models scored below human expert levels overall; with ChatGPT-5.2 Thinking performing closest to experts. Test report interpretation was generally adequate; whereas larger gaps emerged in genetic risk estimation; phenotype prediction; and counseling recommendations (all P < 0.001 vs. human experts; with limited exceptions for ChatGPT-5.2 Thinking in β-thalassemia subtypes). All 145 severe errors were confined to α-thalassemia and α-combined-β-thalassemia subtypes; concentrated in phenotype prediction (76/145; 52.4%) and risk estimation (40/145; 27.6%). LLMs may support knowledge retrieval and structured report interpretation in thalassemia genetic counseling but should not be used independently for risk assessment or final counseling decisions; particularly in complex α-related cases.
Keywords:
thalassemia
large language models
genetic counseling
prenatal diagnosis
performance evaluation

Journal

F
Frontiers in Digital Health
IF:
3.8
Papers:
2.0K
Citations:
3.1K

Organization

D
Department of Obstetrics
Scholars:
595
Papers: 269
Citations: 0
D
Department of Respiratory and Critical Care Medicine
Scholars:
1.2K
Papers: 417
Citations: 0
D
department of clinical laboratory
Scholars:
1.3K
Papers: 471
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers