ArticleFrontiers in public health2026
ChatGPT-5.4 in health education: inter-generation stability and persistent readability challenges.
Article in Frontiers in public health, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
4 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Objectives: Generative artificial intelligence (generative AI) has been investigated for creating patient education materials (PEMs) to reduce the burden of clinical education and improve health information accessibility. However, prior studies have highlighted limitations in readability, content stability, source transparency, and generation quality. The release of ChatGPT-5.4 provides an opportunity to evaluate the inter-generation stability and usability of a newer frontier model. This study evaluates ChatGPT-5.4's performance in producing PEMs for spinal surgery. Method: On 5 March 2026, ChatGPT-5.4 was used to address common patient questions about three prevalent spinal surgeries: lumbar disc herniation surgery, spinal fusion surgery, and spinal decompression surgery. Each question generated five independent responses. Qualitative analysis evaluated language sophistication, information depth, structural clarity, and supplementary content. Readability was measured using the Flesch-Kincaid Reading Ease (FKRE), Flesch-Kincaid Grade Level (FKGL), and Simple Measure of Gobbledygook (SMOG). Quality was assessed by two independent reviewers using the DISCERN tool, with the Intraclass Correlation Coefficient (ICC) calculated. Results: Core medical information remained consistent across generated versions for all three question types, although variations occurred in tone, organization, and detail presentation. The average FKRE ranged from 43.02 to 57.16 ("difficult" to "standard English"), FKGL from 9.12 to 13.34, and SMOG from 8.90 to 11.84, corresponding to reading levels from advanced junior high to early university. The average DISCERN score ranged from 43.2 to 44.9 ("average" quality). ChatGPT-5.4 showed moderate-to-high inter-generation stability with limited content drift in this exemplar context. However because no earlier model was evaluated head-to-head under identical prompts, raters, and time points, this finding should be interpreted as within- study evidence of stability rather than evidence of improved stability over earlier models. Readability remained challenging, and verifiable references were absent. Conclusion: Within this exemplar spinal-surgery context, ChatGPT-5.4 demonstrated moderate-to-high inter-generation consistency in lexical content and core medical themes. Complex language and lack of traceable references may limit accessibility and patient trust. Practice implications: ChatGPT-5.4 may support spinal-surgery patient education by generating stable, clinically plausible PEMs, though factual accuracy was not independently verified. Readability remains above recommended health- literacy levels, and clinician review and plain-language optimization are required before patient use.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.