Evidence map›Paper›PMID 41810293›Full record

ArticleFrontiers in public health2026

Can large language models be trusted? Reliability and readability of responses to perinatal depression FAQs.

Jingyu Huang, Hua Yu, Junjian Chen, Xinyue Wang, Lizhi Huang, Junjie Wen, Hui Li

Erratum issuedAbstract read
In one paragraph

Article in Frontiers in public health, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. An erratum has been issued. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

7 authors.

Jingyu HuangFaculty of Health Sciences, University of Macau, Taipa, China.
Hua YuDepartment of Nursing, Ruikang Hospital Affiliated to Guangxi University of Chinese Medicine, Nanning, China.
Junjian ChenSecond Affiliated Hospital of Guangxi Medical University, Nanning, China.
Xinyue WangGuangXi University of Chinese Medicine, Nanning, China.
Lizhi HuangDepartment of Medical Informatics, Harbin Medical University, Harbin, China.
Junjie WenSouthwest Jiaotong University Hope College, Chengdu, China.
Hui LiDepartment of Nursing, Ruikang Hospital Affiliated to Guangxi University of Chinese Medicine, Nanning, China.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Objective: Large language models (LLMs), a core technology of generative artificial intelligence (AI), are increasingly used in health education and promotion. Although they may expand access to medical information, concerns remain regarding the reliability and readability of AI generated content for the public. This study evaluated the reliability and readability of answers generated by five LLMs to common questions about perinatal depression. The primary aims were to determine (1) the reliability of LLM responses to frequently asked questions about perinatal depression and (2) whether the readability of the generated content aligns with public health literacy levels. Methods: Twenty-seven frequently asked questions were derived from Google Trends and patient facing resources from the American College of Obstetricians and Gynecologists (ACOG). Each question was submitted to ChatGPT-5, Gemini-2.5, Microsoft Copilot, Grok4, and DeepSeek. Two obstetricians independently rated responses using five validated instruments (DISCERN, EQIP, JAMA, GQS, and HONCODE) and inter-rater agreement was quantified using the interclass correlation coefficient (ICC). Readability was assessed using six indices: ARI, GFI, CLI, OLWF, LWGLF, and FRF. Differences among models were analyzed using the Friedman test. Results: Inter rater agreement was high across 27 perinatal depression questions. ICC values ranged from 0.729 to 0.847. Significant between model differences emerged for DISCERN, EQIP, and HONCODE. All had Conclusion: Most LLMs demonstrated moderate to high reliability when responding to perinatal depression questions, supporting their potential as supplementary sources of health information. However, readability levels above recommended benchmarks suggest that current outputs may remain challenging for individuals with lower health literacy. While LLMs improve information accessibility, further improvements in readability, source attribution, and ethical transparency are needed to maximize public benefit and support equitable health communication. Future work should focus on defining and standardizing safety behaviors in high-risk mental health contexts to enable reliable clinical deployment.

Indexed as

ComprehensionDepressionHealth LiteracyLarge Language ModelsFemaleGenerative Artificial IntelligenceHumansPregnancyReproducibility of ResultsSurveys and Questionnairesgenerative artificial intelligencehealth information qualitylarge language modelsperinatal depressionpostpartum depressionreadability

Identifiers

PMID41810293
PMCPMC12968175

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.