Evidence map›Paper›PMID 42498839›Full record

ArticleSkeletal radiology2026

Performance of multimodal large language models versus clinicians for radiographic knee osteoarthritis grading: A multiobserver study.

Asli Irmak Akdogan, Efe Kemal Akdogan, Mehmet Fatih Tumer, Sılanaz Kutlu, Mustafa Agah Tekindal, Ozgur Tosun

Abstract read
PubMed Publisher
In one paragraph

Article in Skeletal radiology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Asli Irmak AkdoganDepartment of Radiology, Izmir Katip Celebi University, Ataturk Training and Research Hospital, Izmir, Turkey. irmakbiranci@gmail.com.ORCID http://orcid.org/0000-0002-6262-1799
Efe Kemal AkdoganDepartment of Orthopedics and Traumatology, Bakircay University Cigli Training and Research Hospital, Izmir, Turkey.
Mehmet Fatih TumerDepartment of Radiology, Izmir Katip Celebi University, Ataturk Training and Research Hospital, Izmir, Turkey.
Sılanaz KutluDepartment of Radiology, Izmir Katip Celebi University, Ataturk Training and Research Hospital, Izmir, Turkey.
Mustafa Agah TekindalDepartment of Biostatistics, Izmir Katip Celebi University, Izmir, Turkey.
Ozgur TosunDepartment of Radiology, Izmir Katip Celebi University, Ataturk Training and Research Hospital, Izmir, Turkey.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

objectiveTo compare the performance of clinicians and two generations of multimodal large language models (LLMs) in Kellgren-Lawrence (KL) grading of knee osteoarthritis (KOA), including feature-level assessment and intraobserver repeatability. MATERIALS AND

methodsIn this retrospective single-center study, 348 knee radiographs were graded by a senior musculoskeletal radiologist (reference standard), a radiologist, an orthopedic surgeon, a radiology resident, and LLMs (ChatGPT-4o and ChatGPT-5.0). Binary KOA detection (KL 0-1 vs ≥ 2), feature-level interpretation (joint space narrowing, osteophytes, subchondral sclerosis), and intraobserver repeatability were evaluated. Agreement metrics included weighted κ, accuracy, and standard diagnostic measures.

resultsAgreement with the reference standard was highest for the radiologist (κ = 0.87), followed by the orthopedic surgeon and radiology resident. Both LLMs demonstrated moderate agreement, with ChatGPT-5.0 outperforming ChatGPT-4o. For binary KOA detection, ChatGPT-5.0 showed very high sensitivity (0.96) but reduced specificity. Per-grade classification was most accurate for KL 0 and KL 4, but remained limited for KL 1-2. Feature-level concordance was modest across all radiographic findings. Intraobserver repeatability was highest for the reference reader (κ = 0.881), followed by the orthopedic surgeon (κ = 0.634) and radiologist (κ = 0.626), while lower agreement was observed for the resident (κ = 0.478) and LLMs, with ChatGPT-5.0 showing higher consistency than ChatGPT-4o (κ = 0.591 vs. 0.485).

conclusionAlthough ChatGPT-5.0 outperformed ChatGPT-4o, both models remained inferior to clinicians in detailed KL grading, feature-level interpretation, and reproducibility. Current multimodal LLMs show high sensitivity but limited specificity and are not suitable for standalone radiographic KOA assessment.

Indexed as

Artificial IntelligenceKnee osteoarthritis; Kellgren-Lawrence gradingLarge Language ModelsRadiography

Identifiers

What Socratic holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.