Evidence map›Paper›PMID 41524886›Full record

ArticleEuropean journal of epidemiology2026

Machine learning versus logistic regression for propensity score estimation: a trial emulation benchmarked against the PARADIGM-HF randomized trial.

Kaicheng Wang, Lindsey Rosman, Haidong Lu

Abstract read
PubMed Publisher
In one paragraph

Article in European journal of epidemiology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

3 authors.

Kaicheng WangYale Center for Analytical Sciences, Yale School of Public Health, New Haven, CT, USA. kaicheng_wang@med.unc.edu.ORCID http://orcid.org/0000-0003-2519-2670
Lindsey RosmanDivision of Cardiology, Department of Medicine, University of North Carolina at Chapel Hill, 160 Dental Circle CB# 7075, Chapel Hill, NC, 27599, USA.
Haidong LuDepartment of Internal Medicine, Yale School of Medicine, New Haven, CT, USA.

Funding

Gender Differences in Physical Activity Among Patients with Implantable Cardioverter Defibrillators (ICDs)K23HL141644 · NHLBI · UNIV OF NORTH CAROLINA CHAPEL HILL · PI ROSMAN, LINDSEY · 2019 to 2023
$910k
Evaluating and Optimizing Care for Opioid Use Disorder using a Structured Data-Science ApproachK99DA057487 · NIDA · YALE UNIVERSITY · PI LU, HAIDONG · 2023 to 2024
$327k
NHLBI NIH HHS K23HL141644NIDA NIH HHS K99DA057487
6 · The paper itself

Abstract

Machine learning (ML) algorithms are increasingly used to estimate propensity score with expectation of improving causal inference. However, the validity of data-driven ML-based approaches for confounder selection and adjustment remains unclear. In this study, we emulated the device-stratified secondary analysis of the PARADIGM-HF trial among U.S. veterans with heart failure and implanted cardiac devices from 2016 to 2020. We benchmarked observational estimates from three propensity score approaches against the trial results. (1) logistic regression with pre-specified confounders (2), generalized boosted models (GBM) using the same pre-specified confounders, and (3) GBM with expanded covariates and automated feature selection. Logistic regression-based propensity score approach yielded estimates closest to the trial (HR = 0.93, 95% CI 0.61-1.42; 23-month RR = 0.86, 95% CI 0.57-1.24 vs. trial HR = 0.81, 95% CI 0.61-1.06). Despite better predictive performance, GBM with pre-specified confounders showed no improvement over the logistic regression approach (HR = 0.97, 95% CI 0.68-1.37; RR = 0.96, 95% CI 0.89-1.98). Moreover, GBM with expanded covariates and data-driven automated feature selection substantially increased bias (HR = 0.61, 95% CI 0.30-1.23; RR = 0.69, 95% CI 0.36-1.04). Our findings suggest that ML-based propensity score methods do not inherently improve causal estimation possibly due to residual confounding from omitted or partially adjusted variables and may introduce overadjustment bias when combined with automated feature selection. These results underscore the importance of careful confounder specification and causal reasoning over algorithmic complexity in causal inference.

Indexed as

Heart FailureMachine LearningPropensity ScoreAgedBoosting Machine Learning AlgorithmsFemaleHumansLogistic ModelsMaleMiddle AgedPrediction AlgorithmsPredictive Learning ModelsUnited StatesBenchmarkingHeart failureMachine learningPropensity score

Identifiers

PMID41524886

What Socratic holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.