ArticleComputational and structural biotechnology journal2025
Evaluating ChatGPT, Gemini and other Large Language Models (LLMs) in orthopaedic diagnostics: A prospective clinical study.
Article in Computational and structural biotechnology journal, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. It is linked to 2 registered trials, which are not on this map. Cited by 19 papers, 1 of them a synthesis that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Observational Study on the Accuracy and Completeness of General Artificial Intelligence in the Diagnosis and Therapeutic Recommendations for Failed or Painful Total Hip Arthroplasty
A Comparative Performance Evaluation of Four Publicly Available Large Language Models Against Gold Standard Medical References
Who cites it
19 citing papers in PubMed, 1 synthesis or guideline pooled it.
- Clinical applications of large language models in knee osteoarthritis: a systematic review.Frontiers in medicine · 2025Pooled it
- Stepwise Diagnostic Evaluation of Chinese Large Language Models: Comparative Study of Common and Rare Diseases.Journal of medical Internet research · 2026Article
- Artificial Intelligence Can Direct Patients Toward a Complaint-specific Musculoskeletal Provider.Journal of the American Academy of Orthopaedic Surgeons. Global research & reviews · 2026Article
- Diagnostic Performance of Large Language Models for Orthopedic-Related Rare Diseases and Their Impact on Physicians' Diagnostic Accuracy: 2-Stage Comparative Evaluation Study Based on the Chinese Rare Disease Catalog.Journal of medical Internet research · 2026Article
- Diagnostic Performance and Workup Efficiency of Large Language Models in Secondary Hypertension: A Blinded Comparative Study.Diagnostics (Basel, Switzerland) · 2026Article
- The doctors of the future: the competition of ChatGPT-4, ChatGPT-4 omni, and Gemini 2.0 Flash in andrology.BMC urology · 2026Article
- Performance of DeepSeek V3.2 and ChatGPT 5.1 in Musculoskeletal Triage and Differential Diagnosis of Outpatients With Low Back Pain: Multidimensional Comparative Study.Journal of medical Internet research · 2026Article
- Promising performance of locally deployed large language models for postoperative orthopaedic patient questions: An In Silico analysis.Journal of experimental orthopaedics · 2026Article
- Diagnostic Performance of Contemporary Large Language Models on Free-Text Histopathologic Descriptions in Oral and Maxillofacial Pathology.Head and neck pathology · 2026Article
- Comprehensive Evaluation of AI Consent Forms in Otolaryngologic Surgery.World journal of otorhinolaryngology - head and neck surgery · 2026Article
- Gemini 1.5 Flash provides the most reliable content while ChatGPT-4o offers the highest readability for patient education on meniscal tears.Knee surgery, sports traumatology, arthroscopy : official journal of the ESSKA · 2026Article
- Artificial Intelligence (AI) in rheumatology: a comparative evaluation of the ChatGPT and DeepSeek application.BMC rheumatology · 2026Article
- Comparative Evaluation of Popular Gen-AI Chatbots in Generating Patient Education Material on Pulmonary Artery Catheter Insertion.Annals of cardiac anaesthesia · 2026Article
- Performance of ChatGPT-4o, Claude 3 Opus, and DeepSeek-R1 in BI-RADS Category 4 Classification and Malignancy Prediction From Mammography Reports: Retrospective Diagnostic Study.JMIR medical informatics · 2025Article
- Vision-based diagnostic gain of ChatGPT-5 and gemini 2.5 pro compared with human experts in oral lesion assessment.Scientific reports · 2025Article
- Artificial intelligence in osteoarthritis research: summary of the 2025 OARSI pre-congress workshop.Osteoarthritis and cartilage open · 2025Article
- Assessing the Clinical Utility of Multimodal Large Language Models in the Diagnosis and Management of Pigmented Choroidal Lesions.Translational vision science & technology · 2025Article
- Article
- Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
9 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Background: Large Language Models (LLMs) such as ChatGPT are gaining attention for their potential applications in healthcare. This study aimed to evaluate the diagnostic sensitivity of various LLMs in detecting hip or knee osteoarthritis (OA) using only patient-reported data collected via a structured questionnaire, without prior medical consultation. Methods: A prospective observational study was conducted at an orthopaedic outpatient clinic specialized in hip and knee OA treatment. A total of 115 patients completed a paper-based questionnaire covering symptoms, medical history, and demographic information. The diagnostic performance of five different LLMs-including four versions of ChatGPT, two of Gemini, Llama, Gemma 2, and Mistral-Nemo-was analysed. Model-generated diagnoses were compared against those provided by experienced orthopaedic clinicians, which served as the reference standard. Results: GPT-4o achieved the highest diagnostic sensitivity at 92.3 %, significantly outperforming other LLMs. The completeness of patient responses to symptom-related questions was the strongest predictor of accuracy for GPT-4o (p < 0.001). Inter-model agreement was moderate among GPT-4 versions, whereas models such as Llama-3.1 demonstrated notably lower accuracy and concordance. Conclusions: GPT-4o demonstrated high accuracy and consistency in diagnosing OA based solely on patient-reported questionnaires, underscoring its potential as a supplementary diagnostic tool in clinical settings. Nevertheless, the reliance on patient-reported data without direct physician involvement highlights the critical need for medical oversight to ensure diagnostic accuracy. Further research is needed to refine LLM capabilities and expand their utility in broader diagnostic applications.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.