Evidence map›Paper›PMID 39850460›Full record

ArticleComputational and structural biotechnology journal2025

Evaluating ChatGPT, Gemini and other Large Language Models (LLMs) in orthopaedic diagnostics: A prospective clinical study.

Stefano Pagano, Luigi Strumolo, Katrin Michalk, Julia Schiegl, Loreto C Pulido, Jan Reinhard, Guenther Maderbacher, Tobias Renkawitz, Marie Schuster

2 registry-linked trialsAbstract read
In one paragraph

Article in Computational and structural biotechnology journal, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. It is linked to 2 registered trials, which are not on this map. Cited by 19 papers, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
19citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

NCT07012577 completednot on this map

Observational Study on the Accuracy and Completeness of General Artificial Intelligence in the Diagnosis and Therapeutic Recommendations for Failed or Painful Total Hip Arthroplasty

TypeobservationalSponsorIstituto Ortopedico RizzoliRan2025 to 2025Enrolled20ConditionsTotal Hip Arthroplasty (THA)ArmsGPT-4 Assessment, Arthroplasty Fellow Assessment, Specializing Resident (4th year) Assessment, Junior Resident (3rd year) Assessment
NCT07199231 active not recruitingnot on this map

A Comparative Performance Evaluation of Four Publicly Available Large Language Models Against Gold Standard Medical References

TypeobservationalSponsorCambridge Health AllianceRan2025 to 2026Enrolled20ConditionsAI (Artificial Intelligence), Large Language Model, Generative Artificial IntelligenceArmsAI clinical reference tool
3 · Its place in the literature

Who cites it

19 citing papers in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
  2. Article
  3. Artificial Intelligence Can Direct Patients Toward a Complaint-specific Musculoskeletal Provider.Journal of the American Academy of Orthopaedic Surgeons. Global research & reviews · 2026
    Article
  4. Article
  5. Article
  6. Article
  7. Article
  8. Article
  9. Article
  10. Comprehensive Evaluation of AI Consent Forms in Otolaryngologic Surgery.World journal of otorhinolaryngology - head and neck surgery · 2026
    Article
  11. Article
  12. Article
  13. Article
  14. Article
  15. Article
  16. Article
  17. Article
  18. Article
  19. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Stefano PaganoDepartment of Orthopaedic Surgery, University of Regensburg, Asklepios Klinikum, Bad Abbach, Germany.
Luigi StrumoloFreelance health consultant & senior data analyst, Avellino, Italy.
Katrin MichalkDepartment of Orthopaedic Surgery, University of Regensburg, Asklepios Klinikum, Bad Abbach, Germany.
Julia SchieglDepartment of Orthopaedic Surgery, University of Regensburg, Asklepios Klinikum, Bad Abbach, Germany.
Loreto C PulidoDepartment of Orthopaedics Hospital of Trauma Surgery, Marktredwitz Hospital, Marktredwitz, Germany.
Jan ReinhardDepartment of Orthopaedic Surgery, University of Regensburg, Asklepios Klinikum, Bad Abbach, Germany.
Guenther MaderbacherDepartment of Orthopaedic Surgery, University of Regensburg, Asklepios Klinikum, Bad Abbach, Germany.
Tobias RenkawitzDepartment of Orthopaedic Surgery, University of Regensburg, Asklepios Klinikum, Bad Abbach, Germany.
Marie SchusterDepartment of Orthopaedic Surgery, University of Regensburg, Asklepios Klinikum, Bad Abbach, Germany.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Large Language Models (LLMs) such as ChatGPT are gaining attention for their potential applications in healthcare. This study aimed to evaluate the diagnostic sensitivity of various LLMs in detecting hip or knee osteoarthritis (OA) using only patient-reported data collected via a structured questionnaire, without prior medical consultation. Methods: A prospective observational study was conducted at an orthopaedic outpatient clinic specialized in hip and knee OA treatment. A total of 115 patients completed a paper-based questionnaire covering symptoms, medical history, and demographic information. The diagnostic performance of five different LLMs-including four versions of ChatGPT, two of Gemini, Llama, Gemma 2, and Mistral-Nemo-was analysed. Model-generated diagnoses were compared against those provided by experienced orthopaedic clinicians, which served as the reference standard. Results: GPT-4o achieved the highest diagnostic sensitivity at 92.3 %, significantly outperforming other LLMs. The completeness of patient responses to symptom-related questions was the strongest predictor of accuracy for GPT-4o (p < 0.001). Inter-model agreement was moderate among GPT-4 versions, whereas models such as Llama-3.1 demonstrated notably lower accuracy and concordance. Conclusions: GPT-4o demonstrated high accuracy and consistency in diagnosing OA based solely on patient-reported questionnaires, underscoring its potential as a supplementary diagnostic tool in clinical settings. Nevertheless, the reliance on patient-reported data without direct physician involvement highlights the critical need for medical oversight to ensure diagnostic accuracy. Further research is needed to refine LLM capabilities and expand their utility in broader diagnostic applications.

Indexed as

Artificial intelligence in healthcareChatGPTDiagnostic sensitivityGeminiGemma 2GPT-4oHip osteoarthritisKnee osteoarthritisLarge Language Models (LLMs)LlamaMistral-NemoMusculoskeletal disordersOrthopaedic diagnosticsPatient-reported data

Identifiers

PMID39850460
PMCPMC11754967

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.