1
Return

ChatGPT provides accurate and safe responses to patient questions on hip arthroscopy, while completeness remains variable: A systematic review and single-arm meta-analysis

delete2026-04-01
delete0
delete
OA
AI
N
Nikolai Ramadanov *
П
Пламен Пенчев
M
Maximilian Voss
H
Heinz, Maximilian
P
Picillo, Marina
M
Mikhail Salzmann
B
Becker, Roland
O
Osterberger, Timoty
B
Banke, Ingo J.
DOI:10.1002/ksa.70396delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Purpose Large language models, such as ChatGPT, are increasingly used by patients seeking information on hip arthroscopy (HAS) and femoroacetabular impingement (FAI). Despite their linguistic fluency, the accuracy, completeness and safety of procedure-specific patient information remain unclear. Although orthopaedic studies report variable performance across subspecialties, no systematic evaluation has specifically addressed HAS. Methods PubMed, Embase, Scopus, CINAHL and Epistemonikos were searched to 10 February 2026 for studies evaluating ChatGPT responses to patient-oriented HAS or FAI questions. Randomized and non-randomized studies, observational cohorts and case series were eligible. Data on question sources, model versions, rating systems and performance domains (accuracy, relevance, completeness, safety, readability and clarity) were extracted. Heterogeneous rating scales were dichotomized into high- versus low-quality responses. Risk of bias was assessed using QUADAS-2 and ROBINS-I. Random-effects single-arm meta-analyses (REML) were conducted for each domain. Results Eight studies met eligibility criteria. Accuracy was high (pooled 88.6%). Relevance, safety, readability and clarity reached pooled values of 100% with low heterogeneity. Completeness was lower (83.8%) with moderate heterogeneity, mainly driven by early GPT-3.5 studies. Funnel plots showed no clear small-study effects, although interpretation was limited by the small number of studies. Risk of bias was predominantly high or moderate, largely due to non-systematic question selection and heterogeneous rating tools. Later models (GPT-4/4o and beyond) demonstrated higher performance compared with GPT-3.5. Conclusion ChatGPT provides accurate, relevant, safe and clear responses to patient questions about HAS, while completeness shows moderate variability. Although LLMs appear promising as adjuncts to patient education, methodological limitations in the current evidence base underscore the need for expert clinical counselling and more rigorous, standardized evaluation frameworks.Level of Evidence Level III, systematic review and meta-analysis of non-randomized studies.
Keywords:
accuracy
ChatGPT
completeness
femoroacetabular impingement
hip arthroscopy
meta-analysis
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Knee Surgery Sports Traumatology Arthroscopy cover
Knee Surgery Sports Traumatology Arthroscopy
IF:
5
Papers:
9.0K
Citations:
2.5W

Organization

Medical University Plovdiv cover
Medical University Plovdiv
Scholars:
1.5K
Papers: 946
Citations: 889
Cited Papers

Cited Papers

Citing Papers

Citing Papers