Evidence map›Paper›PMID 40457713›Full record

ArticlePharmacoepidemiology and drug safety2025

Use of Machine Learning to Compare Disease Risk Scores and Propensity Scores Across Complex Confounding Scenarios: A Simulation Study.

Yuchen Guo, Victoria Y Strauss, Sara Khalid, Daniel Prieto-Alhambra

Abstract readComparative Study
In one paragraph

Article in Pharmacoepidemiology and drug safety, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. Article
  2. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Yuchen GuoCentre for Statistics in Medicine, University of Oxford, Oxford, UK.ORCID https://orcid.org/0000-0002-0847-4855
Victoria Y StraussBoehringer-Ingelheim Pharma GmbH & co., KG, Germany.
Sara KhalidCentre for Statistics in Medicine, University of Oxford, Oxford, UK.
Daniel Prieto-AlhambraCentre for Statistics in Medicine, University of Oxford, Oxford, UK.ORCID https://orcid.org/0000-0002-3950-6346

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

purposeThe surge of treatments for COVID-19 in the second quarter of 2020 had a low prevalence of treatment and high outcome risk. Motivated by that, we conducted a simulation study comparing disease risk scores (DRS) and propensity scores (PS) using a range of scenarios with different treatment prevalences and outcome risks.

methodFour methods were used to estimate PS and DRS: logistic regression (reference method), least absolute shrinkage and selection operator (LASSO), multilayer perceptron (MLP), and XgBoost. Monte Carlo simulations generated data across 25 scenarios varying in treatment prevalence, outcome risk, data complexity, and sample size. Average treatment effects were calculated after matching. Relative bias and average absolute standardized mean difference (ASMD) were reported.

resultEstimation bias increased as treatment prevalence decreased. DRS showed lower bias than PS when treatment prevalence was below 0.1, especially in nonlinear data. However, DRS did not outperform PS in linear or small sample data. PS had comparable or lower bias than DRS when treatment prevalence was 0.1-0.5. Three machine learning (ML) methods performed similarly, with LASSO and XgBoost outperforming the reference method in some nonlinear scenarios. ASMD results indicated that DRS was less impacted by decreasing treatment prevalence compared to PS.

conclusionUnder nonlinear data, DRS reduced bias compared to PS in scenarios with low treatment prevalence, while PS was preferable for data with treatment prevalence greater than 0.1, regardless of the outcome risk. ML methods can outperform the logistic regression method for PS and DRS estimation. Both decreasing sample size and adding nonlinearity and nonadditivity in data increased bias for all methods tested.

Indexed as

COVID-19COVID-19 Drug TreatmentMachine LearningPropensity ScoreBiasComputer SimulationConfounding Factors, EpidemiologicHumansLogistic ModelsMonte Carlo MethodPrevalenceRisk AssessmentSARS-CoV-2causal inferencedisease risk scoresmachine learningpropensity scorestreatment effect

Identifiers

PMID40457713
PMCPMC12130674

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.