Return
Clinical decision support in nasopharyngeal carcinoma: comparative evaluation of large reasoning and language models
DOI:10.1038/s41746-026-03167-3.png)
Abstract
En 中文
This study compared the performance of large language models (LLMs) and large reasoning models (LRMs) in addressing clinical management challenges associated with nasopharyngeal carcinoma (NPC). Five AI models, three LLMs (ChatGPT-4, ChatGPT-4o, and Gemini 2.0 Flash) and two LRMs (Deepseek-R1 and Grok 3 [Think]), were evaluated using 50 custom-designed open-ended questions spanning five NPC management modules. Responses were independently scored by two radiation oncologists in a single-blind manner. Overall, Grok 3 (Think) and Deepseek-R1 achieved higher scores than ChatGPT-4 and Gemini 2.0 Flash, and showed higher proportions of superior responses in complex domains such as Multidisciplinary Treatment and Radiotherapy. In multidimensional assessment, Grok 3 (Think) achieved the highest accuracy (84.0%) and relevance (91.6%), whereas Deepseek-R1 excelled in comprehensiveness (83.2%). However, all models exhibited notable limitations, including outdated content, hallucinations, and limited evidence traceability. LRMs tended to outperform LLMs in responding to open-ended clinical questions on NPC management. The difference reached statistical significance in selected overall pairwise comparisons and certain dimension-specific metrics, with descriptive advantages observed in module-specific analyses. These findings suggest that LRMs have potential utility as assistive tools for clinical decision support. However, rigorous validation and cautious interpretation of AI-generated content remain essential to ensure reliability in clinical practice.
Journal
IF:
15.1
Papers:
3.1K
Citations:
1.5W

