Evidence mapPaperPMID 41316099Full record

ArticleBMC medical informatics and decision making2025

Augmenting small tabular health data for training prognostic ensemble machine learning models using generative models.

Dan Liu, Samer El Kababji, Nicholas Mitsakakis, Lisa Pilgram, Thomas D Walters, Mark Clemons, Gregory R Pond, Alaa El-Hussuna, Khaled El Emam

Abstract read
In one paragraph

Article in BMC medical informatics and decision making, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Dan LiuChildren's Hospital of Eastern Ontario Research Institute, 401 Smyth Road, Ottawa, ON, K1h 8l1, Canada.
Samer El KababjiChildren's Hospital of Eastern Ontario Research Institute, 401 Smyth Road, Ottawa, ON, K1h 8l1, Canada.
Nicholas MitsakakisChildren's Hospital of Eastern Ontario Research Institute, 401 Smyth Road, Ottawa, ON, K1h 8l1, Canada.
Lisa PilgramChildren's Hospital of Eastern Ontario Research Institute, 401 Smyth Road, Ottawa, ON, K1h 8l1, Canada.
Thomas D WaltersHospital for Sick Children, Toronto, ON, Canada.
Mark ClemonsOttawa Hospital Research Institute, Ottawa, ON, Canada.
Gregory R PondMcMaster University, Hamilton, ON, Canada.
Alaa El-HussunaOpenSourceResearch, Aalborg, Denmark.
Khaled El EmamChildren's Hospital of Eastern Ontario Research Institute, 401 Smyth Road, Ottawa, ON, K1h 8l1, Canada. kelemam@ehealthinformation.ca.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundSmall datasets are common in health research. However, the generalization performance of machine learning models is suboptimal when the training datasets are small. To address this, data augmentation is one solution and is often used for imaging and time series data, but there are no evaluations on its potential benefits for tabular health data. Augmentation increases sample size and is seen as a form of regularization that increases the diversity of small datasets, leading them to perform better on unseen data.

objectivesEvaluate data augmentation using generative models on tabular health data and assess the impact of diversity versus increasing the sample size.

methodsUsing 13 large health datasets, we performed a simulation to evaluate the impact of data augmentation on the prediction performance (as measured by the ROC-AUC, the area under the receiver operating characteristic curve) on binary classification gradient boosted decision tree models. Four different synthetic data generation models were evaluated. We also built a generalized linear mixed effect model to assess the variable importance for model performance improvements from augmentation. We illustrate the proposed method on seven small real datasets as an application. A comparison of augmentation with resampling (which is a proxy for a larger dataset with minimal impact on diversity) was performed.

resultsAugmentation improves prognostic performance for datasets that have higher cardinality categorical variables and lower baseline ROC-AUC. No specific generative model consistently outperformed the others. For the seven small application datasets, augmenting the existing data results in an increase in ROC-AUC between 4.31% (ROC-AUC from 0.71 to 0.75) and 43.23% (ROC-AUC from 0.51 to 0.73), with an average 15.55% relative improvement, demonstrating the nontrivial impact of augmentation on small datasets (p = 0.0078). Augmentation ROC-AUC was higher than resampling only ROC-AUC (p = 0.016). The diversity of augmented datasets was higher than the diversity of resampled datasets (p = 0.046).

conclusionsThis study demonstrates that data augmentation using generative models can have a marked benefit in terms of improved predictive performance for machine learning models on tabular health data, but only for datasets that meet baseline data complexity and predictive performance criteria. Our mixed effect model identified the most influential characteristics of the dataset and can help end-users have a more realistic expectation of the augmentation performance for a new dataset. Furthermore, augmentation performed better when having a smaller dataset, which is consistent with the argument that greater data diversity due to augmentation is beneficial. CLINICAL

trial registrationNot applicable.

Indexed as

Datasets as TopicMachine LearningBoosting Machine Learning AlgorithmsClassification AlgorithmsData AnalyticsGenerative Artificial IntelligenceHumansPrediction AlgorithmsPredictive Learning ModelsPrognosisArtificial intelligenceData augmentationData scarcityGenerative modelsMachine learningSynthetic data

Identifiers

PMID41316099
PMCPMC12661835

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.