Return
Performance and limitations of four large language models in genetic counseling for thalassemia
W
J
M
Q
Y
J
J
秦
Y
Y
DOI:10.3389/fdgth.2026.1902480.png)
Abstract
En 中文
In regions with high thalassemia prevalence; such as southern China and Southeast Asia; chronic shortages of professional genetic counseling resources have driven interest in large language models (LLMs) as auxiliary tools; yet their performance and safety boundaries in this setting remain uncharacterized. This single-center retrospective study evaluated four LLMs (ChatGPT-5.2 Thinking; DeepSeek-V3.2 chat; Gemini 3 Flash; and Grok 4.1 Fast) using 1; 080 standardized knowledge questions administered across five independent sessions and 150 real-world clinical cases scored by six senior experts across eight dimensions. All models exceeded 90% accuracy on single-choice and true-false questions. Between-model differences were most pronounced in multiple-choice questions; where ChatGPT-5.2 Thinking achieved the highest accuracy (87.28% ± 1.69%); significantly outperforming Grok 4.1 Fast (72.06% ± 1.69%; P < 0.001). DeepSeek-V3.2 chat showed lower cross-session consistency than the other three models. In clinical case analysis; all models scored below human expert levels overall; with ChatGPT-5.2 Thinking performing closest to experts. Test report interpretation was generally adequate; whereas larger gaps emerged in genetic risk estimation; phenotype prediction; and counseling recommendations (all P < 0.001 vs. human experts; with limited exceptions for ChatGPT-5.2 Thinking in β-thalassemia subtypes). All 145 severe errors were confined to α-thalassemia and α-combined-β-thalassemia subtypes; concentrated in phenotype prediction (76/145; 52.4%) and risk estimation (40/145; 27.6%). LLMs may support knowledge retrieval and structured report interpretation in thalassemia genetic counseling but should not be used independently for risk assessment or final counseling decisions; particularly in complex α-related cases.
Keywords:
thalassemia
large language models
genetic counseling
prenatal diagnosis
performance evaluation
Journal
F
IF:
3.8
Papers:
2.0K
Citations:
3.1K
