Evidence map›Paper›PMID 42110661›Full record

ReviewJournal of healthcare informatics research2026

A Scoping Review of Synthetic Data Generation by Language Models in Biomedical Research and Application: Data Utility and Quality Perspectives.

Hanshu Rao, Weisi Liu, Haohan Wang, I-Chan Huang, Zhe He, Xiaolei Huang

Abstract readReview
In one paragraph

Review in Journal of healthcare informatics research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
  2. Article
  3. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Hanshu RaoDepartment of Computer Science, University of Memphis, Memphis, 38152 TN United States.
Weisi LiuDepartment of Computer Science, University of Memphis, Memphis, 38152 TN United States.
Haohan WangSchool of Information Sciences, University of Illinois Urbana-Champaign, Champaign, 61820 IL United States.
I-Chan HuangEpidemiology and Cancer Control, St Jude Children's Research Hospital, Memphis, 38105 TN United States.
Zhe HeSchool of Information, Florida State University, Tallahassee, 32306 FL United States.
Xiaolei HuangDepartment of Computer Science, University of Memphis, Memphis, 38152 TN United States.

Funding

Patient-Generated Health Data to Predict Childhood Cancer Survivorship OutcomesR01CA258193 · NCI · ST. JUDE CHILDREN'S RESEARCH HOSPITAL · PI I-Chan Huang, YUTAKA YASUI · 2021 to 2026
$3.7M
The Adherence Promotion with Person-centered Technology (APPT) Project: Promoting Adherence to Enhance the Early Detection and Treatment of Cognitive DeclineR01AG064529 · NIA · FLORIDA STATE UNIVERSITY · PI BOOT, WALTER RICHARD, CHAKRABORTY, SHAYOK · 2019 to 2023
$3.2M
Prediction of Health Outcomes and Adverse Events in Pediatric Organ Transplantation in FloridaR21LM013911 · NLM · FLORIDA STATE UNIVERSITY · PI HE, ZHE, KILLIAN, MICHAEL · 2022 to 2023
$395k
Precision HIV Prevention: Piloting a youth learning health communityR21MH137736 · NIMH · FLORIDA STATE UNIVERSITY · PI HE, ZHE, NAAR, SYLVIE · 2024 to 2025
$346k
Toward Deep Learning Techniques for Cell-Type and Spatial Resolution Estimation of Regulatory NetworksR03OD038389 · OD · UNIVERSITY OF ILLINOIS AT URBANA-CHAMPAIGN · PI WANG, HAOHAN · 2024 to 2024
$290k
AHRQ HHS R21 HS029969NCI NIH HHS R01 CA258193NIA NIH HHS R01 AG064529NIH HHS R03 OD038389NIMH NIH HHS R21 MH137736NLM NIH HHS R21 LM013911
6 · The paper itself

Abstract

Synthetic data generation using large language models (LLMs) demonstrates substantial promise in addressing biomedical data challenges and shows increasing adoption in biomedical research. This study systematically reviews recent advances in synthetic data generation for biomedical applications and clinical research, focusing on how LLMs address data scarcity, utility, and quality issues with different modalities. We conducted a scoping review following PRISMA-ScR guidelines and searched literature published between 2020 and 2025 through PubMed, ACM, Web of Science, and Google Scholar. A total of 59 studies were included based on relevance to synthetic data generation in biomedical contexts. Among the reviewed studies, the predominant data modalities were unstructured texts (78.0%), tabular data (13.6%), and multimodal sources (8.4%). Common generation methods included LLM prompting (74.6%), fine-tuning (20.3%), and specialized models (5.1%). Evaluations were heterogeneous: intrinsic metrics (27.1%), human-in-the-loop assessments (44.1%), and LLM-based evaluations (13.6%). However, limitations and key barriers persist in data modalities, domain utility, resource and model accessibility, and standardized evaluation protocols. Future efforts may focus on developing standardized, transparent evaluation frameworks and expanding accessibility to support effective applications in biomedical research. Supplementary Information: The online version contains supplementary material available at 10.1007/s41666-026-00229-9.

Indexed as

Biomedical informaticsData qualityLanguage modelSynthetic data generation

Identifiers

PMID42110661
PMCPMC13156355

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.