Evidence map›Paper›PMID 41707656›Full record

ArticleCell reports. Medicine2026

Benchmarking large language models for predictive modeling in biomedical research with a focus on reproductive health.

Reuben Sarwal, Victor Tarca, Claire A Dubin, Nikolas Kalavros, Gaurav Bhatti, Sanchita Bhattacharya, Atul Butte, Roberto Romero, Gustavo Stolovitzky, Tomiko T Oskotsky and 2 more

Abstract read
In one paragraph

Article in Cell reports. Medicine, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

12 authors.

Reuben SarwalBakar Computational Health Sciences Institute, University of California, San Francisco, San Francisco, CA 94158, USA.
Victor TarcaHuron High School, Ann Arbor, MI 48105, USA.
Claire A DubinBakar Computational Health Sciences Institute, University of California, San Francisco, San Francisco, CA 94158, USA.
Nikolas KalavrosNew York University Langone Health, New York, NY 10016, USA; Division of Precision Medicine, Department of Medicine, NYU Grossman School of Medicine, New York, NY 10016, USA.
Gaurav BhattiCenter for Molecular Medicine and Genetics, Wayne State University School of Medicine, Detroit, MI 48201, USA.
Sanchita BhattacharyaBakar Computational Health Sciences Institute, University of California, San Francisco, San Francisco, CA 94158, USA.
Atul ButteBakar Computational Health Sciences Institute, University of California, San Francisco, San Francisco, CA 94158, USA; Department of Pediatrics, University of California, San Francisco, San Francisco, CA 94143, USA.
Roberto RomeroPregnancy Research Branch, Division of Obstetrics and Maternal-Fetal Medicine, Division of Intramural Research, Eunice Kennedy Shriver National Institute of Child Health and Human Development, National Institutes of Health, United States Department of Health and Human Services, Bethesda, MD 20892, USA; Department of Obstetrics and Gynecology, University of Michigan, Ann Arbor, MI 48109, USA; Department of Epidemiology and Biostatistics, Michigan State University, East Lansing, MI 48824, USA.
Gustavo StolovitzkyDepartment of Pathology, New York University Grossman School of Medicine, New York, NY 10016, USA; Biomedical Data Science Hub, New York University Langone Health, New York, NY 10016, USA.
Tomiko T OskotskyBakar Computational Health Sciences Institute, University of California, San Francisco, San Francisco, CA 94158, USA; Division of Clinical Informatics and Digital Transformation, Department of Medicine, University of California, San Francisco, San Francisco, CA 94143, USA.
Adi L TarcaCenter for Molecular Medicine and Genetics, Wayne State University School of Medicine, Detroit, MI 48201, USA; Department of Obstetrics and Gynecology, Wayne State University School of Medicine, Detroit, MI 48201, USA; Department of Computer Science, Wayne State University College of Engineering, Detroit, MI 48201, USA.
Marina SirotaBakar Computational Health Sciences Institute, University of California, San Francisco, San Francisco, CA 94158, USA; Department of Pediatrics, University of California, San Francisco, San Francisco, CA 94143, USA. Electronic address: marina.sirota@ucsf.edu.

Funding

Placenta-specific maternal plasma proteomic biomarkers of fetal deathR21HD115800 · NICHD · WAYNE STATE UNIVERSITY · PI TARCA, ADI LAURENTIU · 2025 to 2025
$424k
NICHD NIH HHS R21 HD115800
6 · The paper itself

Abstract

Large language models (LLMs) are increasingly used for code generation and data analysis. This study assesses LLM performance across four predictive tasks from three DREAM challenges: gestational age regression from transcriptomics and DNA methylation and classification of preterm birth and early preterm birth from microbiome data. We prompt LLMs with task descriptions, data locations, and target outcomes and then run LLM-generated code to fit prediction models and determine accuracy on test sets. Among the eight LLMs tested, o3-mini-high, 4o, DeepseekR1, and Gemini 2.0 can complete at least one task. R code generation is more successful (14/16) than Python (7/16). OpenAI's o3-mini-high outperforms others, completing 7/8 tasks. Test set performance of the top LLM-generated models matches or exceeds the median-participating team for all four tasks and surpasses the top-performing team for one task (p = 0.02). These findings underscore the potential of LLMs to democratize predictive modeling in omics and increase research output.

Indexed as

BenchmarkingBiomedical ResearchLarge Language ModelsReproductive HealthDNA MethylationFemaleGestational AgeHumansPredictive Learning ModelsPregnancyPremature BirthbenchmarkingDREAM challengeslarge language modelsomics dataplacenta clockpredictive analyticspreterm birthreproductive health

Identifiers

PMID41707656
PMCPMC12923944

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.