Evidence map›Paper›PMID 39756334›Full record

ArticleSurgery2025

De novo generation of colorectal patient educational materials using large language models: Prompt engineering key to improved readability.

India E Ellison, Wendelyn M Oslock, Abiha Abdullah, Lauren Wood, Mohanraj Thirumalai, Nathan English, Bayley A Jones, Robert Hollis, Michael Rubyan, Daniel I Chu

Abstract readComparative Study
In one paragraph

Article in Surgery, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Article
  2. Article
  3. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

India E EllisonDepartment of Surgery, University of Alabama at Birmingham, AL.
Wendelyn M OslockDepartment of Surgery, University of Alabama at Birmingham, AL; Department of Quality, Birmingham Veterans Affairs Medical Center, AL. Electronic address: https://www.twitter.com/WendelynOslock.
Abiha AbdullahTrauma and Transfusion Department, University of Pittsburgh Medical College, PA. Electronic address: https://www.twitter.com/abihaabdullah7.
Lauren WoodDepartment of Surgery, University of Alabama at Birmingham, AL.
Mohanraj ThirumalaiDepartment of Surgery, University of Alabama at Birmingham, AL.
Nathan EnglishDepartment of Surgery, University of Alabama at Birmingham, AL; Department of General Surgery, University of Cape Town, WC, South Africa.
Bayley A JonesDepartment of Surgery, University of Alabama at Birmingham, AL; Department of Surgery, University of Texas Southwestern Medical Center, Dallas, TX. Electronic address: https://www.twitter.com/bayley_jones.
Robert HollisDepartment of Surgery, University of Alabama at Birmingham, AL. Electronic address: https://www.twitter.com/rhhollis.
Michael RubyanUniversity of Michigan School of Public Health, Ann Arbor, MI.
Daniel I ChuDepartment of Surgery, University of Alabama at Birmingham, AL. Electronic address: dchu@uabmc.edu.

Funding

Designing a Plan of Action for Better Access and Quality of Surgery for African-Americans with Gastrointestinal Cancers in the Deep SouthR01MD013858 · NIMHD · UNIVERSITY OF ALABAMA AT BIRMINGHAM · PI CHU, DANIEL I, PISU, MARIA · 2020 to 2024
$3.0M
Adapting Enhanced Recovery Programs (ERPs) through Health Literacy to Eliminate Surgical DisparitiesR01CA271303 · NCI · UNIVERSITY OF ALABAMA AT BIRMINGHAM · PI Daniel I Chu · 2023 to 2026
$2.1M
Enhancing Health Literacy in Surgery (EHLIS) to Eliminate Surgical Disparities for African-Americans with Inflammatory Bowel Disease (IBD)K23MD013903 · NIMHD · UNIVERSITY OF ALABAMA AT BIRMINGHAM · PI CHU, DANIEL I · 2019 to 2021
$434k
AHRQ HHS K12 HS023009NCI NIH HHS R01 CA271303NIMHD NIH HHS K23 MD013903NIMHD NIH HHS R01 MD013858
6 · The paper itself

Abstract

backgroundImproving patient education has been shown to improve clinical outcomes and reduce disparities, though such efforts can be labor intensive. Large language models may serve as an accessible method to improve patient educational material. The aim of this study was to compare readability between existing educational materials and those generated by large language models.

methodsBaseline colorectal surgery educational materials were gathered from a large academic institution (n = 52). Three prompts were entered into Perplexity and ChatGPT 3.5 for each topic: a Basic prompt that simply requested patient educational information the topic, an Iterative prompt that repeated instruction asking for the information to be more health literate, and a Metric-based prompt that requested a sixth-grade reading level, short sentences, and short words. Flesch-Kincaid Grade Level or Grade Level, Flesch-Kincaid Reading Ease or Ease, and Modified Grade Level scores were calculated for all materials, and unpaired t tests were used to compare mean scores between baseline and documents generated by artificial intelligence platforms.

resultsOverall existing materials were longer than materials generated by the large language models across categories and prompts: 863-956 words vs 170-265 (ChatGPT) and 220-313 (Perplexity), all P < .01. Baseline materials did not meet sixth-grade readability guidelines based on grade level (Grade Level 7.0-9.8 and Modified Grade Level 9.6-11.5) or ease of readability (Ease 53.1-65.0). Readability of materials generated by a large language model varied by prompt and platform. Overall, ChatGPT materials were more readable than baseline materials with the Metric-based prompt: Grade Level 5.2 vs 8.1, Modified Grade Level 7.3 vs 10.3, and Ease 70.5 vs 60.4, all P < .01. In contrast, Perplexity-generated materials were significantly less readable except for those generated with the Metric-based prompt, which did not statistically differ.

conclusionBoth existing materials and the majority of educational materials created by large language models did not meet readability recommendations. The exception to this was with ChatGPT materials generated with a Metric-based prompt that consistently improved readability scores from baseline and met recommendations in terms of the average Grade Level score. The variability in performance highlights the importance of the prompt used with large language models.

Indexed as

Colorectal SurgeryComprehensionHealth LiteracyPatient Education as TopicTeaching MaterialsHumansLarge Language Models

Identifiers

PMID39756334
PMCPMC11936715

What Socratic holds

Textmetadata
LicenceTDM
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.