Return
When Large Language Models (LLMs) walk into a Bachelor's in optometry examination: Comparing the performances of LLMs and bachelor of optometry students
DOI:10.4103/IJO.IJO_808_25.png)
Abstract
En 中文
Purpose:To evaluate the performance of Large Language Models (LLMs) on optometry examination questions and compare their accuracy and readability with Bachelor of Optometry students.Methods:A cross-sectional comparative study was conducted using the publicly available, free versions of five LLM models from four platforms (ChatGPT 3.5, ChatGPT 4o, Gemini, CoPilot, and DeepSeek) and a group of 15 third- and fourth-year optometry students. Two sets of multiple-choice questions (20 theoretical and 20 clinical) were administered to both the students and the LLMs. Theoretical questions covered core optometric knowledge, while clinical questions simulated real-life patient scenarios. Responses were graded by senior ophthalmologists for accuracy, and readability was assessed via readable.com using four indices, including Flesch-Kincaid Grade Level, Flesch Reading Ease Score, Coleman Liau Score, and Simple Measure of Gobbledygook (SMOG) Index.Results:The overall scores of the optometry students (28.13 +/- 3.33) were comparable to those of the LLMs (29 +/- 4.41). In theoretical questions, LLMs (15.40 +/- 1.82) performed at par with the students (14.07 +/- 2.21), with DeepSeek and CoPilot outperforming students (scoring 17 each). However, in clinical questions, the students performed better, highlighting the limitations of LLMs in context-specific reasoning. Pairwise comparisons of the readability analysis revealed that Gemini and DeepSeek provided significantly most readable explanations, while ChatGPT 3.5 produced the most complex responses. Across models, readability varied for Flesch-Kincaid grade level (P = 0.0213), Flesch Reading Ease Score (P = 0.0014), and SMOG (P = 0.0412), with a nonsignificant trend for Coleman Liau Score (P = 0.0529).Conclusion:LLMs show reasonable accuracy, matching students in theoretical performance but underperforming in clinical reasoning. Gemini and DeepSeek offer superior readability, highlighting their promise as educational tools. Future research should focus on integrating LLMs into curricula while balancing them with hands-on clinical education.
Keywords:
Artificial Intelligence
ChatGPT
clinical reasoning
DeepSeek
Gemini
large language models
medical education
optometry
readability
Journal
I
IF:
1.8
Papers:
220
Citations:
9.0K

