ArticleAmerican journal of medical genetics. Part A2025
Diagnostic Accuracy of a Custom Large Language Model on Rare Pediatric Disease Case Reports.
Article in American journal of medical genetics. Part A, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 18 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
18 citing papers in PubMed.
- Attitudes Toward Large Language Models in Health Care and Preferences for Their Adoption and Oversight Among Health Care Professionals: Cross-Sectional Survey.Journal of medical Internet research · 2026Article
- Machine Learning, Large Language Models, and Multimodal AI for Diagnosing Pediatric Rare Diseases: Scoping Review.Journal of medical Internet research · 2026Article
- Disparate language and model effects on AI-based translation and recognition of genetic conditions.Journal of the American Medical Informatics Association : JAMIA · 2026Article
- Beyond accuracy metrics: Toward responsible integration of large language models in pediatric diagnostic reasoning.Pediatric investigation · 2026Article
- Large-language-models for pediatric diagnosis: Performance evaluation using real-world clinical notes from common and rare cases.Pediatric investigation · 2026Article
- Performance of ChatGPT-4o in Providing Information on Pediatric Inborn Errors of Immunity: A Cross-Sectional Evaluation.Journal of clinical medicine · 2026Article
- Large Language Models for Diagnosis and Prognosis of Chronic Liver Diseases: A Systematic Review.Health science reports · 2026Review
- Reframing AI for Rare Disease Recognition.Research square · 2026Article
- Large Language Model Performance and Clinical Reasoning Tasks.JAMA network open · 2026Article
- GEN-KnowRD: Reframing AI for Rare Disease Recognition.medRxiv : the preprint server for health sciences · 2026Article
- Designing Patient-Centered Communication Aids in Pediatric Surgery Using Large Language Models.Journal of pediatric surgery · 2025Article
- Development and Evaluation of an Artificial Intelligence-Powered Surgical Oral Examination Simulator: A Pilot Study.Mayo Clinic proceedings. Digital health · 2025Article
- Improving automated deep phenotyping through large language models using retrieval-augmented generation.Genome medicine · 2025Article
- Large Language Models in Medical Diagnostics: Scoping Review With Bibliometric Analysis.Journal of medical Internet research · 2025Article
- Synthetic medical education in dermatology leveraging generative artificial intelligence.NPJ digital medicine · 2025Article
- A Future of Self-Directed Patient Internet Research: Large Language Model-Based Tools Versus Standard Search Engines.Annals of biomedical engineering · 2025Article
- Artificial intelligence in clinical genetics.European journal of human genetics : EJHG · 2025Review
- Reliability of ChatGPT answers to common questions on developmental dysplasia of the hip as an information source for parents.Frontiers in pediatrics · 2025Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
7 authors.
Funding
Abstract
Accurately diagnosing rare pediatric diseases frequently represent a clinical challenge due to their complex and unusual clinical presentations. Here, we explore the capabilities of three large language models (LLMs), GPT-4, Gemini Pro, and a custom-built LLM (GPT-4 integrated with the Human Phenotype Ontology [GPT-4 HPO]), by evaluating their diagnostic performance on 61 rare pediatric disease case reports. The performance of the LLMs were assessed for accuracy in identifying specific diagnoses, listing the correct diagnosis among a differential list, and broad disease categories. In addition, GPT-4 HPO was tested on 100 general pediatrics case reports previously assessed on other LLMs to further validate its performance. The results indicated that GPT-4 was able to predict the correct diagnosis with a diagnostic accuracy of 13.1%, whereas both GPT-4 HPO and Gemini Pro had diagnostic accuracies of 8.2%. Further, GPT-4 HPO showed an improved performance compared with the other two LLMs in identifying the correct diagnosis among its differential list and the broad disease category. Although these findings underscore the potential of LLMs for diagnostic support, particularly when enhanced with domain-specific ontologies, they also stress the need for further improvement prior to integration into clinical practice.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.