Evidence map›Paper›PMID 40825542›Full record

ArticleJournal of medical Internet research2025

Magnitude and Impact of Hallucinations in Tabular Synthetic Health Data on Prognostic Machine Learning Models: Validation Study.

Lisa Pilgram, Samer El Kababji, Dan Liu, Khaled El Emam

Abstract readValidation Study
In one paragraph

Article in Journal of medical Internet research, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Lisa PilgramSchool of Epidemiology and Public Health, Faculty of Medicine, University of Ottawa, Ottawa, ON, Canada.ORCID https://orcid.org/0000-0002-1020-0650
Samer El KababjiCHEO Research Institute, Children's Hospital of Eastern Ontario, Ottawa, ON, Canada.ORCID https://orcid.org/0000-0002-7642-2280
Dan LiuSchool of Epidemiology and Public Health, Faculty of Medicine, University of Ottawa, Ottawa, ON, Canada.ORCID https://orcid.org/0000-0001-9632-4736
Khaled El EmamSchool of Epidemiology and Public Health, Faculty of Medicine, University of Ottawa, Ottawa, ON, Canada.ORCID https://orcid.org/0000-0003-3325-4149

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundGenerative artificial intelligence (AI) for tabular synthetic data generation (SDG) has significant potential to accelerate health care research and innovation. A critical limitation of generative AI, however, is hallucinations. Although this has been commonly observed in text-generating models, it may also occur in tabular SDG.

objectiveThis study aims to investigate the magnitude of hallucinations in tabular synthetic data, whether their frequency increases with training data complexity, and the extent to which they impact the utility of synthetic data for downstream prognostic machine learning (ML) modeling tasks.

methodsOn the basis of 12 large and high-dimensional real-world health care datasets, 6354 training datasets of different complexity were created by varying the subset of variables included in each dataset. Synthetic data were generated using 7 different SDG models. Hallucinations were defined as synthetic records that did not exist in the population, and the hallucination rate (HR) was the proportion of hallucinations in a synthetic dataset. Classification was the downstream prognostic modeling task, conducted via an ML approach (light gradient boosted machine) and an artificial neural network (multilayer perceptron). Mixed-effects models were fitted to examine the relationship between training data complexity and the HR and the HR and the predictive performance of AI and ML models when trained on the synthetic data.

resultsThe HR ranged from 0.3% to 100% (median 99.1%, IQR 98.5%-100.0%) and increased with training data complexity. However, in most SDG models, the HR did not affect AI and ML prognostic model performance. In the SDG models in which a significant association was detected, the estimated effect was very small, with a maximum decrease in the area under the receiver operating characteristic curve of -0.0002 (95% CI -0.0003 to -0.0002, P<.001) in light gradient boosting machine and -0.0001 (95% CI -0.0002 to -0.0001, P=.002) in multilayer perceptron.

conclusionsThese findings suggest that while hallucinations may be very common in synthetic tabular health data, they do not necessarily impair its utility for prognostic modeling.

Indexed as

HallucinationsMachine LearningArtificial IntelligenceHumansNeural Networks, ComputerPrognosisAIartificial intelligencedata utilitygenerative modelshallucinationssynthetic data

Identifiers

PMID40825542
PMCPMC12402739

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.