Evidence mapPaperPMID 42404600Full record

ArticleDigital health

Challenges of patient-facing generative artificial intelligence in hypertension care: A cross-platform evaluation of the quality, readability, and actionability of LLM-Generated patient education materials.

Mengqiu Hu, Zhiqiang Wang, Zhiwen Zhang, Muwei Li

Abstract read
In one paragraph

Article in Digital health. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Mengqiu HuDepartment of Cardiology, Huanggang Central Hospital, Huanggang, Hubei, China.
Zhiqiang WangSchool of Medicine, Yangtze University, Jingzhou, Hubei, China.ORCID https://orcid.org/0009-0004-1036-3638
Zhiwen ZhangDepartment of Cardiology, Fuwai Central China Cardiovascular Hospital (Central China Fuwai Hospital of Zhengzhou University), Zhengzhou, Henan, China.
Muwei LiDepartment of Cardiology, Fuwai Central China Cardiovascular Hospital (Central China Fuwai Hospital of Zhengzhou University), Zhengzhou, Henan, China.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Objective: To assess the quality, readability, and actionability of hypertension patient education materials generated by six selected patient-facing large language model (LLM) platforms, and to characterize cross-platform heterogeneity to guide the optimization of AI-generated patient education materials. Methods: Six publicly accessible or commonly accessible platforms were evaluated using standardized prompts to generate patient education materials. Understandability and actionability were assessed using the Patient Education Materials Assessment Tool for Printable Materials (PEMAT-P). Information quality was evaluated using the expanded Ensuring Quality Information for Patients scale (EQIP-36), and overall quality was rated using the Global Quality Score (GQS). Readability was compared using seven metrics, including the Flesch Reading Ease Score (FRES) and the Flesch-Kincaid Grade Level (FKGL). Results: Overall educational quality and understandability were generally favorable, but substantial cross-platform heterogeneity was observed. Readability remained challenging, with low FRES values and grade-level indices generally exceeding commonly recommended thresholds for patient education materials. Qwen3-Max-Thinking-Preview achieved the highest PEMAT-P total score (77.00), followed by ChatGPT 5.2-Thinking (72.22). For EQIP-36, Qwen3-Max-Thinking-Preview scored highest (49.28), followed by DeepSeek-R1 (45.48). DeepSeek-R1 generated the most readable materials among the evaluated platforms, with a median FRES of approximately 42.14 and FKGL of approximately 10.10, whereas Qwen3-Max-Thinking-Preview showed lower readability, with a median FRES of approximately 16.82 and FKGL of approximately 14.38. Kimi K2 showed high PEMAT-P understandability (76.90) but low actionability (10.00). Post hoc analyses showed that Qwen3-Max-Thinking-Preview significantly outperformed ERNIE Bot 4.5 Turbo, Doubao, and Kimi K2 on PEMAT-P total score, and outperformed Kimi K2 and ERNIE Bot 4.5 Turbo on EQIP-36. DeepSeek-R1 also outperformed Kimi K2 and ERNIE Bot 4.5 Turbo on EQIP-36. Across content domains, actionability was significantly higher for Daily Care and Prevention than for Basic Understanding of the Disease and Complications, Psychological and Social Aspects. Conclusions: Generative AI shows promise for hypertension patient education, particularly in improving understandability. However, actionability remains a major limitation of current outputs, highlighting the need for platform-aware optimization strategies that explicitly strengthen step-by-step and action-oriented guidance.

Indexed as

actionabilitycross-platform evaluationgenerative artificial intelligencehypertensionlarge language modelspatient educationunderstandability

Identifiers

PMID42404600
PMCPMC13332286

What Socratic holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.