Evidence mapPaperPMID 42525673Full record

ArticlePloS one2026

A comparative analysis of readability, quality, and reliability in large language model outputs pertaining to knee osteoarthritis queries.

Erdem Maraşlı, Erkan Ozduran, Volkan Hancı

Abstract readComparative Study
In one paragraph

Article in PloS one, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

3 authors.

Erdem MaraşlıAydın State Hospital, Physical Medicine and Rehabilitation, Pain Medicine, Aydin, Turkey.
Erkan OzduranSivas Numune Hospital, Physical Medicine and Rehabilitation, Pain Medicine, Sivas, Turkey.ORCID https://orcid.org/0000-0003-3425-313X
Volkan HancıDokuz Eylul University, Anesthesiology and Reanimation, Critical Care Medicine, Izmir, Turkey.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

This study aims to comparatively examine the readability, accuracy, and quality of responses provided by artificial intelligence (AI)-based chatbots such as Perplexity, ChatGPT-5, and Gemini to questions about knee osteoarthritis (KOA), which accounts for approximately four-fifths of the global osteoarthritis (OA) burden. In this study, 8 keywords were determined by excluding repetitive, irrelevant or synonymous ones from the 25 most frequently used English keywords associated with KOA based on Google Trends data, and these terms were asked as questions to three different artificial intelligence-based chatbots. The study measured readability using formulas like Coleman-Liau Index (CLI), Automated Readability Index (ARI), and Linsear Write (LW). Reliability of the information was assessed using the Journal of the American Medical Association (JAMA) benchmarks along with the modified DISCERN instrument. To determine overall content quality, the Global Quality Score (GQS) and the Ensuring Quality Information for Patients (EQIP) scale were applied. Together, these tools provided a comprehensive assessment of how understandable, reliable, and high-quality each chatbot's responses were. The most frequently searched keywords related to OA were "osteoarthritis of knee," "knee pain," and "osteoarthritis knee pain." A readability analysis of responses from three different AI-based chat systems revealed that all platforms had text levels above the Grade 6 threshold, and this difference was statistically significant (p < 0.05). Comparisons demonstrated that ChatGPT-5 produced the most readable content (FRES:45, GFOG:11.9, FKGL:9.24, CLI:14.03, SMOG:8.37, ARI:11.37, LW:7.2). However, Perplexity achieved significantly higher scores than ChatGPT-5 across all quality and reliability assessments, yielding superior median scores (DISCERN: 4, JAMA: 2, GQS: 4, EQIP: 92.8). Perplexity also outperformed Gemini in the mDISCERN reliability assessment (p = 0.001), while no significant difference in quality or reliability was found between Gemini and ChatGPT-5. No statistically significant difference was found between Gemini and ChatGPT in reliability and quality surveys. This analysis of KOA highlights significant challenges regarding the potential of popular AI chatbots for patient information. When examining readability levels, responses from these tools consistently exceed the recommended comprehensibility threshold, making it difficult for patients to absorb critical information. Furthermore, the relatively low scores recorded in reliability and content quality assessments raise significant concerns about the scientific validity and integrity of the medical information presented. Given these findings, the sufficient quality, robustness, and appropriate levels of understandability of future AI-based tools can only be ensured by the establishment and operation of an effective oversight mechanism.

Indexed as

ComprehensionOsteoarthritis, KneeArtificial IntelligenceHumansLarge Language ModelsReproducibility of Results

Identifiers

PMID42525673
PMCPMC13419221

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.