Evidence map›Paper›PMID 41625978›Full record

ArticleFrontiers in bioengineering and biotechnology2026

Benchmarking readability, reliability, and scientific quality of large language models in communicating organoid science.

Man Sun, Dan Zang, Jun Chen

Abstract read
In one paragraph

Article in Frontiers in bioengineering and biotechnology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

3 authors.

Man Sun *Department of Oncology, The Second Hospital of Dalian Medical University, Dalian, Liaoning, China.
Dan Zang *Department of Oncology, The Second Hospital of Dalian Medical University, Dalian, Liaoning, China.
Jun ChenDepartment of Oncology, The Second Hospital of Dalian Medical University, Dalian, Liaoning, China.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Organoids have become central platforms in precision oncology and translational research, increasing the need for communication that is accurate, transparent, and clinically responsible. Large language models (LLMs) are now widely consulted for organoid-related explanations, but their ability to balance readability, scientific rigor, and educational suitability has not been systematically established. Methods: Five mainstream LLMs (GPT-5, DeepSeek, Doubao, Tongyi Qianwen, and Wenxin Yiyan) were systematically evaluated using a curated set of thirty representative organoid-related questions. For each model, twenty outputs were independently scored using the C-PEMAT-P scale, the Global Quality Score (GQS), and seven validated readability indices. Between-model differences were analyzed using one-way ANOVA or Kruskal-Wallis tests, and correlation analyses were performed to examine associations between readability and quality measures. Results: Model performance differed markedly, with GPT-5 achieving the highest C-PEMAT and GQS scores (16.05 ± 1.10; 4.70 ± 0.47; both P < 0.001), followed by intermediate performance from DeepSeek and Doubao (C-PEMAT 11.75 ± 2.07 and 12.05 ± 1.82; GQS 3.65 ± 0.49 and 3.35 ± 0.49). Tongyi Qianwen and Wenxin Yiyan comprised the lowest-performing tier (C-PEMAT 7.85 ± 1.09 and 9.00 ± 2.05; GQS 1.55 ± 0.51 and 2.10 ± 0.55). Score-distribution patterns further highlighted reliability gaps, with GPT-5 showing tightly clustered values and domestic models displaying broader dispersion and unstable performance. Readability differed significantly across models and question categories, with safety-related, diagnostic, and technical questions showing the highest linguistic and conceptual complexity. Correlation analyses showed strong internal coherence among readability indices but only weak-to-moderate associations with C-PEMAT, GQS, and reliability metrics, indicating that linguistic simplicity is not a dependable surrogate for scientific quality. Conclusion: LLMs exhibited substantial variability in communicating organoid-related information, forming distinct performance tiers with direct implications for patient education and translational decision-making. Because readability, scientific quality, and reliability diverged across models, linguistic simplification alone is insufficient to guarantee accurate or dependable interpretation. These findings underscore the need for organoid-adapted AI systems that integrate domain-specific knowledge, convey uncertainty transparently, ensure output reliability, and safeguard safety-critical information.

Indexed as

artificial intelligencelarge language modelsonline medical informationorganoidsreadability

Identifiers

PMID41625978
PMCPMC12855483

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.