Evidence map›Paper›PMID 41735989›Full record

ArticleBMC medical informatics and decision making2026

Simulated evaluation of large language model stepwise diagnostic reasoning with real-world chest pain encounters and Bayesian networks.

Conrad W Safranek, Vimig Socrates, Donald Wright, Thomas Huang, Alaa Alashi, Kent McCann, R Andrew Taylor, David Chartash

Abstract read
In one paragraph

Article in BMC medical informatics and decision making, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Conrad W Safranek *Department of Biomedical Informatics and Data Science, School of Medicine, Yale University, New Haven, CT, USA.
Vimig Socrates *Department of Biomedical Informatics and Data Science, School of Medicine, Yale University, New Haven, CT, USA.
Donald WrightDepartment of Biomedical Informatics and Data Science, School of Medicine, Yale University, New Haven, CT, USA.
Thomas HuangDepartment of Emergency Medicine, School of Medicine, Yale University, New Haven, CT, USA.
Alaa AlashiDepartment of Internal Medicine, Cardiovascular Medicine Section, Yale University, New Haven, CT, USA.
Kent McCannDepartment of Emergency Medicine, School of Medicine, Yale University, New Haven, CT, USA.
R Andrew TaylorDepartment of Biomedical Informatics and Data Science, School of Medicine, Yale University, New Haven, CT, USA. KQC5MK@uvahealth.org.
David ChartashDepartment of Emergency Medicine, School of Medicine, Yale University, New Haven, CT, USA. david.chartash@yale.edu.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundReal-world evaluation of large language models (LLMs) as clinical diagnostic aids is limited by the reliance on static vignettes and retrospective data, which inadequately reflect the dynamic, iterative nature of clinical decision-making and may overestimate LLMs' performance. Here, we benchmark GPT-4o in a stepwise simulated diagnostic setting with real-world clinical data, comparing its diagnostic accuracy and information-seeking strategy with Bayesian-network-derived optimal policies and observed physician practice.

methodsWe assessed GPT-4o across 500 emergency department (ED) chest-pain encounters, drawn from a cohort of 202,632 cases spanning three EDs. A Bayesian network (BN) trained on the structured cohort data imputed clinical data not collected in the original encounter to create a more robust simulation environment. The BN furthermore enabled derivation of mutual-information-optimal query pathways. GPT-4o sequentially requested information from 136 structured clinical variables under three prompting regimes that varied in disease-prevalence cues and diagnostic category constraints. Diagnostic decisions encompassed one of seven predefined emergent conditions or Other Diagnosis. We measured diagnostic accuracy under each prompting strategy, as well as calculated rank-based overlap with the BN optimal pathway to benchmark the LLM's information-seeking behavior.

resultsAcross the full chest-pain cohort, life-threatening etiologies accounted for only 2.14% of encounters (from 1.04% acute coronary syndrome to 0.01% esophageal rupture). With baseline prompting, GPT-4o systematically over-predicted rare conditions (sensitivity 79.3%; specificity 45.2%); adding prevalence cues or removing diagnostic category constraints respectively increased specificity (83.0% and 94.7%) while reducing false alarms by 107 and 140 per 500, but at the cost of poor sensitivity (30.4% and 8.8%). Rank-biased overlap between GPT-4o's information-seeking sequence and the Bayesian-network mutual-information optimum was low across diagnoses (range 0.060-0.097), and the model diverged from clinician behavior by requesting fewer vitals ([Formula: see text]-fold) and labs ([Formula: see text]-fold), while requesting 30%+ more imaging data.

conclusionsIn this simulated assessment, GPT-4o demonstrated diagnostic biases toward rare conditions and differed substantially from normative probabilistic models and physician practice patterns. These discrepancies could lead to unnecessary over-triage and resource utilization. Integrating LLMs within more rigorous probabilistic frameworks and calibrating them to realistic disease prevalences may be essential for effectively harnessing their potential as clinical decision-support tools.

Indexed as

Chest PainClinical Decision-MakingClinical ReasoningLarge Language ModelsBayes TheoremEmergency Service, HospitalFemaleHumansMaleMiddle AgedArtificial intelligenceBayesian networksDecision-support toolsDiagnostic decision-makingDiagnostic pathway optimizationEmergency department diagnosticsLarge Language Models (LLMs)Non-traumatic chest painProbabilistic imputation

Identifiers

PMID41735989
PMCPMC13037032

What Socratic holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.