Evidence map›Paper›PMID 39268988›Full record

ArticleAmerican journal of medical genetics. Part A2025

Diagnostic Accuracy of a Custom Large Language Model on Rare Pediatric Disease Case Reports.

Cameron C Young, Ellie Enichen, Christian Rivera, Corinne A Auger, Nathan Grant, Arya Rao, Marc D Succi

Abstract read
In one paragraph

Article in American journal of medical genetics. Part A, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 18 papers.

0numbers the graph read from it
0cells of the map it votes in
18citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

18 citing papers in PubMed.

  1. Article
  2. Article
  3. Disparate language and model effects on AI-based translation and recognition of genetic conditions.Journal of the American Medical Informatics Association : JAMIA · 2026
    Article
  4. Article
  5. Article
  6. Article
  7. Review
  8. Article
  9. Article
  10. GEN-KnowRD: Reframing AI for Rare Disease Recognition.medRxiv : the preprint server for health sciences · 2026
    Article
  11. Article
  12. Article
  13. Article
  14. Article
  15. Article
  16. Article
  17. Artificial intelligence in clinical genetics.European journal of human genetics : EJHG · 2025
    Review
  18. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Cameron C YoungHarvard Medical School, Boston, Massachusetts, USA.ORCID 0000-0001-7323-4834
Ellie EnichenHarvard Medical School, Boston, Massachusetts, USA.
Christian RiveraHarvard Medical School, Boston, Massachusetts, USA.
Corinne A AugerHarvard Medical School, Boston, Massachusetts, USA.
Nathan GrantHarvard Medical School, Boston, Massachusetts, USA.
Arya RaoHarvard Medical School, Boston, Massachusetts, USA.
Marc D SucciMedically Engineered Solutions in Healthcare Incubator, Innovation in Operations Research Center, Mass General Brigham, Boston, Massachusetts, USA.ORCID 0000-0002-1518-3984

Funding

Medical Scientist Training ProgramT32GM144273 · NIGMS · HARVARD MEDICAL SCHOOL · PI David Shumway Jones, Jacqueline A. Lees · 2022 to 2026
$14.7M
NIGMS NIH HHS T32 GM144273
6 · The paper itself

Abstract

Accurately diagnosing rare pediatric diseases frequently represent a clinical challenge due to their complex and unusual clinical presentations. Here, we explore the capabilities of three large language models (LLMs), GPT-4, Gemini Pro, and a custom-built LLM (GPT-4 integrated with the Human Phenotype Ontology [GPT-4 HPO]), by evaluating their diagnostic performance on 61 rare pediatric disease case reports. The performance of the LLMs were assessed for accuracy in identifying specific diagnoses, listing the correct diagnosis among a differential list, and broad disease categories. In addition, GPT-4 HPO was tested on 100 general pediatrics case reports previously assessed on other LLMs to further validate its performance. The results indicated that GPT-4 was able to predict the correct diagnosis with a diagnostic accuracy of 13.1%, whereas both GPT-4 HPO and Gemini Pro had diagnostic accuracies of 8.2%. Further, GPT-4 HPO showed an improved performance compared with the other two LLMs in identifying the correct diagnosis among its differential list and the broad disease category. Although these findings underscore the potential of LLMs for diagnostic support, particularly when enhanced with domain-specific ontologies, they also stress the need for further improvement prior to integration into clinical practice.

Indexed as

Rare DiseasesChildHumansPediatricsPhenotypeartificial intelligencediagnostic supportgeneticslarge language modelspediatric rare disease

Identifiers

PMID39268988
PMCPMC12123583

What Socratic holds

Textmetadata
LicenceTDM
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.