ArticleJournal of cancer research and clinical oncology2025
Performance of large language models in reporting oral health concerns and side effects in head and neck cancer: a comparative study.
Article in Journal of cancer research and clinical oncology, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
5 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
purposeWith increasing reliance on large language models (LLMs) for health information, this study evaluated reliability and quality, understandability, actionability, readability and misinformation risk of responses from LLMs to oral health concerns and oral side effects in head and neck cancer (HNC) patients.
methodsFrequently asked questions on oral health and HNC therapy side effects were identified via ChatGPT-GPT-4-turbo and Gemini-2.5 Flash, then submitted to eight LLMs (ChatGPT-GPT-4-turbo, Gemini-2.5 Flash, Microsoft Copilot, Perplexity, Chatsonic, Mistral, Meta AI-Llama 4, DeepSeek-R1). Responses were assessed using DISCERN and modified DISCERN instruments (reliability and quality), Patient Education Materials Assessment Tool (PEMAT [understandability and actionability]), Flesch-Reading-Ease-Score (FRES [readability]), misinformation score, citations, and wordcounts. Statistical analysis was done by Scheirer-Ray-Hare-test followed by Dunn's post-hoc-tests and Bonferroni-Holm correction (p < 0.05).
resultsA total of 40 questions belonging to 12 oral health-related categories were identified. Statistically significant differences between LLMs were found for DISCERN, modified DISCERN, PEMAT-understandability, PEMAT-actionability, FRES, and word counts (p < 0.001). Median DISCERN and modified DISCERN scores amounted from 47.0 (ChatGPT-GPT-4-turbo) to 59.0 (Perplexity, Chatsonic) and from 2.0 (Gemini-2.5 Flash, Mistral) to 5.0 (Perplexity) indicating good to fair reliability. LLMs were understandable (median PEMAT-understandability scores ≥ 75.0), but provided limited specific guidance (median PEMAT-actionability scores ≤ 40) and used complex language (median FRES ≤ 40.2). Misinformation risk was generally low and not statistically significant among LLMs (p = 0.768).
conclusionDespite a low overall misinformation risk, deficits in actionability highlight the need for cautious integration of LLMs into HNC patient education.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.