Evidence map›Paper›PMID 40709446›Full record

ArticleStroke2025

Accuracy of Large Language Models to Identify Stroke Subtypes Within Unstructured Electronic Health Record Data.

Dylan Owens, Danh Q Nguyen, Michael Dohopolski, Justin F Rousseau, Eric D Peterson, Ann Marie Navar

Abstract read
In one paragraph

Article in Stroke, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 7 papers.

0numbers the graph read from it
0cells of the map it votes in
7citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

7 citing papers in PubMed.

  1. Observational
  2. Article
  3. Article
  4. Article
  5. Article
  6. Article
  7. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Dylan OwensDepartment of Medicine (D.O., D.Q.N., E.D.P., A.M.N.), UT Southwestern Medical Center, Dallas, TX.ORCID 0009-0009-9883-6001
Danh Q NguyenDepartment of Medicine (D.O., D.Q.N., E.D.P., A.M.N.), UT Southwestern Medical Center, Dallas, TX.ORCID 0000-0003-2451-3583
Michael DohopolskiDepartment of Radiation Oncology (M.D.), UT Southwestern Medical Center, Dallas, TX.ORCID 0000-0002-9043-1490
Justin F RousseauBiostatistics and Clinical Informatics Section, Department of Neurology (J.F.R.), UT Southwestern Medical Center, Dallas, TX.ORCID 0000-0002-2817-9124
Eric D Peterson *Department of Medicine (D.O., D.Q.N., E.D.P., A.M.N.), UT Southwestern Medical Center, Dallas, TX.ORCID 0000-0002-5415-4721
Ann Marie Navar *Department of Medicine (D.O., D.Q.N., E.D.P., A.M.N.), UT Southwestern Medical Center, Dallas, TX.ORCID 0000-0002-6197-9860

Funding

UT Southwestern Center for Translational MedicineUL1TR003163 · NCATS · UT SOUTHWESTERN MEDICAL CENTER · PI TOTO, ROBERT DANIEL · 2021 to 2025
$39.3M
Training in Cardiovascular ResearchT32HL125247 · NHLBI · UT SOUTHWESTERN MEDICAL CENTER · PI JOSEPH A HILL · 2015 to 2026
$4.3M
NCATS NIH HHS UL1 TR003163NHLBI NIH HHS L30 HL181920NHLBI NIH HHS T32 HL125247
6 · The paper itself

Abstract

backgroundWhile

methodsWe implemented a retrieval-augmented generation framework with GPT-4o to classify stroke types (ischemic versus hemorrhagic) and ischemic stroke subtypes using electronic health records data. The American Heart Association Get With The Guidelines-Stroke registry served as the gold standard. Model development used a 20% subset of Get With The Guidelines-Stroke-linked data from UT Southwestern Medical Center (UTSW), with the remaining 80% reserved for testing. External validation used data from the Parkland Health and Hospital System (PHHS). A total of 4123 stroke hospitalizations from January 2019 to August 2023 were included (UTSW: n=2047; PHHS: n=2076). Three prompting strategies-zero-shot chain-of-thought, expert-guided, and instruction-based-were evaluated. Predictions of GPT-4os were compared with classifications made by trained abstractors contributing to the Get With The Guidelines-Stroke registry.

resultsIn the external validation set, 79.6% of patients had ischemic stroke and 20.4% hemorrhagic. GPT-4o achieved 98% accuracy (95% CI, 0.97-0.99) in classifying stroke type, where accuracy reflects the overall proportion of correctly classified patients. Sensitivity was 0.98 (95% CI, 0.97-0.99), and specificity was 0.97 (95% CI, 0.96-0.98). For ischemic stroke subtypes, sensitivity ranged from 0.40 (95% CI, 0.31-0.49) for cryptogenic to 0.95 (95% CI, 0.93-0.97) for small-vessel occlusion. Specificity ranged from 0.94 (95% CI, 0.92-0.96) for large-artery atherosclerosis to 0.98 (95% CI, 0.97-0.99) for cardioembolism. Zero-shot chain-of-thought prompting-requiring minimal human input-performed comparably to more labor-intensive strategies. Consistency analysis revealed

conclusionsGPT-4o demonstrated strong accuracy in classifying stroke types but faced challenges with ischemic subtypes.

Indexed as

Electronic Health RecordsIschemic StrokeLanguageStrokeFemaleHumansLarge Language ModelsMaleRegistriesartificial intelligenceelectronic health recordshemorrhagic strokeInternational Classification of Diseasesischemic stroke

Identifiers

PMID40709446
PMCPMC12313299

What Socratic holds

Textmetadata
LicenceTDM
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.