Evidence map›Paper›PMID 41420748›Full record

ArticleJournal of cancer research and clinical oncology2025

Performance of large language models in reporting oral health concerns and side effects in head and neck cancer: a comparative study.

Jonas Rast, Susanne Wiegand, Jana Biermann, Annette Wiegand, Felix Marschner

Abstract readComparative Study
In one paragraph

Article in Journal of cancer research and clinical oncology, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Jonas RastDepartment of Otorhinolaryngology, Head and Neck Surgery, University Hospital Schleswig-Holstein, Arnold-Heller-Straße 3, 24105, Kiel, Germany.ORCID http://orcid.org/0009-0008-8369-156X
Susanne WiegandDepartment of Otorhinolaryngology, Head and Neck Surgery, University Hospital Schleswig-Holstein, Arnold-Heller-Straße 3, 24105, Kiel, Germany.ORCID http://orcid.org/0000-0003-1183-9226
Jana BiermannDepartment of Preventive Dentistry, Periodontology and Cariology, University Medical Center Göttingen, Robert-Koch-Str. 40, 37075, Göttingen, Germany.ORCID http://orcid.org/0000-0002-0492-7677
Annette WiegandDepartment of Preventive Dentistry, Periodontology and Cariology, University Medical Center Göttingen, Robert-Koch-Str. 40, 37075, Göttingen, Germany.ORCID http://orcid.org/0000-0001-5640-8775
Felix MarschnerDepartment of Preventive Dentistry, Periodontology and Cariology, University Medical Center Göttingen, Robert-Koch-Str. 40, 37075, Göttingen, Germany. felix.marschner@med.uni-goettingen.de.ORCID http://orcid.org/0009-0009-6913-5855

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

purposeWith increasing reliance on large language models (LLMs) for health information, this study evaluated reliability and quality, understandability, actionability, readability and misinformation risk of responses from LLMs to oral health concerns and oral side effects in head and neck cancer (HNC) patients.

methodsFrequently asked questions on oral health and HNC therapy side effects were identified via ChatGPT-GPT-4-turbo and Gemini-2.5 Flash, then submitted to eight LLMs (ChatGPT-GPT-4-turbo, Gemini-2.5 Flash, Microsoft Copilot, Perplexity, Chatsonic, Mistral, Meta AI-Llama 4, DeepSeek-R1). Responses were assessed using DISCERN and modified DISCERN instruments (reliability and quality), Patient Education Materials Assessment Tool (PEMAT [understandability and actionability]), Flesch-Reading-Ease-Score (FRES [readability]), misinformation score, citations, and wordcounts. Statistical analysis was done by Scheirer-Ray-Hare-test followed by Dunn's post-hoc-tests and Bonferroni-Holm correction (p < 0.05).

resultsA total of 40 questions belonging to 12 oral health-related categories were identified. Statistically significant differences between LLMs were found for DISCERN, modified DISCERN, PEMAT-understandability, PEMAT-actionability, FRES, and word counts (p < 0.001). Median DISCERN and modified DISCERN scores amounted from 47.0 (ChatGPT-GPT-4-turbo) to 59.0 (Perplexity, Chatsonic) and from 2.0 (Gemini-2.5 Flash, Mistral) to 5.0 (Perplexity) indicating good to fair reliability. LLMs were understandable (median PEMAT-understandability scores ≥ 75.0), but provided limited specific guidance (median PEMAT-actionability scores ≤ 40) and used complex language (median FRES ≤ 40.2). Misinformation risk was generally low and not statistically significant among LLMs (p = 0.768).

conclusionDespite a low overall misinformation risk, deficits in actionability highlight the need for cautious integration of LLMs into HNC patient education.

Indexed as

Head and Neck NeoplasmsLanguageOral HealthComprehensionFemaleHumansLarge Language ModelsMaleMiddle AgedPatient Education as TopicSurveys and QuestionnairesArtificial intelligenceHead and neck cancerInformation qualityLarge language modelOral healthReadability

Identifiers

PMID41420748
PMCPMC12718290

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.