ArticleKnee surgery, sports traumatology, arthroscopy : official journal of the ESSKA2025
Evaluating DeepResearch and DeepThink in anterior cruciate ligament surgery patient education: ChatGPT-4o excels in comprehensiveness, DeepSeek R1 leads in clarity and readability of orthopaedic information.
Article in Knee surgery, sports traumatology, arthroscopy : official journal of the ESSKA, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. It is linked to trial NCT07631585 (Longitudinal Pre-Post Patient AI Trust Dynamics in Orthopedic Outpatients), which is not on this map. Cited by 23 papers, 1 of them a synthesis that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Longitudinal Pre-Post Patient AI Trust Dynamics in Orthopedic Outpatients: A Mixed-Methods Observational Study With Matched Physician-Patient Dyads
Who cites it
23 citing papers in PubMed, 1 synthesis or guideline pooled it.
- Applications of natural language processing and large language models in sports injury assessment and rehabilitation decision-making: a scoping review.Frontiers in medicine · 2026Pooled it
- DeepSeek-assisted problem-based learning for glaucoma education in an undergraduate ophthalmology clerkship: a randomized educational pilot study.Scientific reports · 2026Trial
- Evaluation of large language model responses to patient questions on oral anticoagulant therapy: a comparative expert assessment.Exploratory research in clinical and social pharmacy · 2026Article
- Automated data extraction for systematic reviews using GPT-5.2 and Google Gemini Pro 3: A dual-large language model approach in orthopaedic research.Knee surgery, sports traumatology, arthroscopy : official journal of the ESSKA · 2026Article
- Comparative performance of AI models and clinicians in evidence-based cardiovascular disease management for people living with HIV: Comparative Study.Journal of medical Internet research · 2026Article
- Article
- Applications, Challenges, and Future Directions of Large Language Models in Health Care Communication: Scoping Review.Journal of medical Internet research · 2026Article
- Applications of DeepSeek in Medicine: Bibliometric Analysis and Scoping Review.Journal of medical Internet research · 2026Article
- Artificial intelligence and machine learning in sports medicine: mapping clinical tasks and assessing clinical maturity - a scoping review.BMC medical informatics and decision making · 2026Article
- Evaluating deepresearch and deepthink in total knee arthroplasty patient education: ChatGPT-4o excels in comprehensiveness, Deepseek R1 leads in clarity and readability of orthopedic information.Joint diseases and related surgery · 2026Article
- Deep Research Agents: Major Breakthrough or Incremental Progress for Medical AI?Journal of medical Internet research · 2026Review
- Gemini 1.5 Flash provides the most reliable content while ChatGPT-4o offers the highest readability for patient education on meniscal tears.Knee surgery, sports traumatology, arthroscopy : official journal of the ESSKA · 2026Article
- Reasoning-optimised large language models reach near-expert accuracy on board-style orthopaedic exams: A multi-model comparison on 702 multiple-choice questions.Knee surgery, sports traumatology, arthroscopy : official journal of the ESSKA · 2026Article
- ChatGPT models provide higher-quality but lower-readability responses than Google Gemini regarding anterior shoulder instability, with no added benefit of the orthopaedic expert plugin.Knee surgery, sports traumatology, arthroscopy : official journal of the ESSKA · 2026Article
- An evaluation of DeepSeek and healthcare professionals' Q&A capabilities in improving patient-family satisfaction in the ICU.Frontiers in medicine · 2026Article
- High patient and surgeon satisfaction with ChatGPT-generated responses to real patient questions regarding total knee arthroplasty.Journal of orthopaedic surgery and research · 2025Article
- Promoting Responsible DeepSeek Deployment in Health Care: Scoping Review Comparing Grey and White Literature.Journal of medical Internet research · 2025Article
- Can Artificial Intelligence Educate Patients? Comparative Analysis of ChatGPT and DeepSeek Models in Meniscus Injuries.Healthcare (Basel, Switzerland) · 2025Article
- ChatGPT provides high-quality answers to FAQs about high tibial osteotomy despite low inter-rater agreement.Journal of experimental orthopaedics · 2025Article
- Artificial intelligence algorithms in orthopaedics: A narrative review of methods and clinical applications.Journal of experimental orthopaedics · 2025Review
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
8 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
purposeThis study compares ChatGPT-4o, equipped with its deep research feature, and DeepSeek R1, equipped with its deepthink feature-both enabling real-time online data access-in generating responses to frequently asked questions (FAQs) about anterior cruciate ligament (ACL) surgery. The aim is to evaluate and compare their performance in terms of accuracy, clarity, completeness, consistency and readibility for evidence-based patient education.
methodsA list of ten FAQs about ACL surgery was compiled after reviewing the Sports Medicine Fellowship Institution's webpages. These questions were posed to ChatGPT and DeepSeek in research-enabled modes. Orthopaedic sports surgeons evaluated the responses for accuracy, clarity, completeness, and consistency using a 4-point Likert scale. Inter-rater reliability of the evaluations was assessed using intraclass correlation coefficients (ICCs). In addition, a readability analysis was conducted using the Flesch-Kincaid Grade Level (FKGL) and Flesch Reading Ease Score (FRES) metrics via an established online calculator to objectively measure textual complexity. Paired t tests were used to compare the mean scores of the two models for each criterion, with significance set at p < 0.05.
resultsBoth models demonstrated high accuracy (mean scores of 3.9/4) and consistency (4/4). Significant differences were observed in clarity and completeness: ChatGPT provided more comprehensive responses (mean completeness 4.0 vs. 3.2, p < 0.001), while DeepSeek's answers were clearer and more accessible to laypersons (mean clarity 3.9 vs. 3.0, p < 0.001). DeepSeek had lower FKGL (8.9 vs. 14.2, p < 0.001) and higher FRES (61.3 vs. 32.7, p < 0.001), indicating greater ease of reading for a general audience. ICC analysis indicated substantial inter-rater agreement (composite ICC = 0.80).
conclusionChatGPT-4o, leveraging its deep research feature, and DeepSeek R1, utilizing its deepthink feature, both deliver high-quality, accurate information for ACL surgery patient education. While ChatGPT excels in comprehensiveness, DeepSeek outperforms in clarity and readability, suggesting that integrating the strengths of both models could optimize patient education outcomes. LEVEL OF EVIDENCE: Level V.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.