Evidence mapPaperPMID 42396348Full record

ArticlemedRxiv : the preprint server for health sciences2026

Development and External Validation of a Machine Learning Model for 10-Year Ischemic Stroke Risk Prediction in Diverse Populations.

Ahmed Khattab, Zhe Wang, Vinodh Srinivasasainagendra, Hemant K Tiwari, Ruth Loos, Nita Limdi, Marguerite Ryan Irvin

Abstract readPreprint
In one paragraph

Article in medRxiv : the preprint server for health sciences, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Ahmed KhattabDepartment of Epidemiology, School of Public Health, The University of Alabama at Birmingham, Birmingham, AL, USA.ORCID 0000-0002-7253-199X
Zhe WangDepartment of Epidemiology, School of Public Health, The University of Alabama at Birmingham, Birmingham, AL, USA.
Vinodh SrinivasasainagendraDepartment of Biostatistics, School of Public Health, The University of Alabama at Birmingham, Birmingham, AL, USA.
Hemant K TiwariDepartment of Biostatistics, School of Public Health, The University of Alabama at Birmingham, Birmingham, AL, USA.
Ruth LoosThe Charles Bronfman Institute for Personalized Medicine, The Icahn School of Medicine at Mount Sinai, NY, USA.
Nita LimdiDepartment of Neurology, Heersink School of Medicine, The University of Alabama at Birmingham, AL, USA.
Marguerite Ryan IrvinDepartment of Epidemiology, School of Public Health, The University of Alabama at Birmingham, Birmingham, AL, USA.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Importance: Machine-learning models for ischemic stroke risk prediction are rarely validated across ancestrally distinct cohorts, and the contributions of polygenic risk scores (PRS) and self-reported race in such models remain unclear. Objective: To develop and externally validate a 10-year ischemic stroke risk model and quantify the incremental contributions of laboratory trajectories, PRS, and self-reported race and ethnicity across populations. Design Setting and Participants: Retrospective cohort study with model development in the All of Us (AoU) Research Program (n = 34,987; 1,920 incident strokes) and external validation in the Bio Exposures: Three XGBoost model tiers added laboratory feature trajectories (M2) and 20 PRS (M3) to clinical baseline features (M1); evaluated under race-blind and race-aware specifications. Main Outcomes and Measures: First inpatient ischemic stroke within 10 years; discrimination (area under the receiver operating characteristic curve [AUROC]) and calibration (observed-to-expected [O/E] ratio). Results: In the AoU test partition (n = 6,998; 384 cases), M3 achieved AUROC 0.813 (95% CI, 0.788-0.837), outperforming the Revised Framingham Stroke Risk Profile (ΔAUROC 0.164) and Pooled Cohort Equations (ΔAUROC 0.181; both Conclusions and Relevance: A machine-learning ensemble combining clinical, laboratory, and polygenic features outperformed traditional risk scores by 0.16-0.18 AUROC and retained discriminative validity in an ancestrally distinct external cohort but required site-specific recalibration of absolute risk. The marginal contribution of self-reported race overlapped with polygenic signal, supporting per-ancestry calibration over universal race-aware model deployment.

Identifiers

PMID42396348
PMCPMC13321198

What Socratic holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.