Evidence map›Paper›PMID 36403757›Full record

ArticleJournal of biomedical informatics2023

COMMUTE: Communication-efficient transfer learning for multi-site risk prediction.

Tian Gu, Phil H Lee, Rui Duan

Abstract read
In one paragraph

Article in Journal of biomedical informatics, 2023. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 14 papers.

0numbers the graph read from it
0cells of the map it votes in
14citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

14 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
  4. Article
  5. Article
  6. Article
  7. On the Connections Among Three Transfer Learning Paradigms.Stat (International Statistical Institute) · 2025
    Article
  8. Article
  9. Robust angle-based transfer learning in high dimensions.Journal of the Royal Statistical Society. Series B, Statistical methodology · 2025
    Article
  10. Article
  11. Review
  12. Review
  13. Article
  14. Multi-Task Learning with Summary Statistics.Advances in neural information processing systems · 2023
    Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

3 authors.

Tian GuDepartment of Biostatistics, Harvard T.H. Chan School of Public Health, Boston, MA, United States.
Phil H LeeDepartment of Psychiatry, Harvard Medical School, Boston, MA, United States; Center for Genomic Medicine, Massachusetts General Hospital, Boston, MA, United States; Stanley Center for Psychiatric Research, Broad Institute of MIT and Harvard, Cambridge, MA, United States.
Rui DuanDepartment of Biostatistics, Harvard T.H. Chan School of Public Health, Boston, MA, United States. Electronic address: rduan@hsph.harvard.edu.

Funding

Discoveries in ADHD genomics: Help or hype in clinical settings?R01MH116037 · NIMH · MASSACHUSETTS GENERAL HOSPITAL · PI DOYLE, ALYSA E · 2019 to 2023
$4.4M
Alignment of cortical development trajectories with emergent dimensional psychopathology and related risk factors among early adolescents in the ABCD StudyR01MH124694 · NIMH · MASSACHUSETTS GENERAL HOSPITAL · PI Joshua Lawrence Roffman · 2020 to 2026
$2.9M
Comprehensive analysis of genetic pleiotropy in eleven neuropsychiatric disordersR01MH119243 · NIMH · MASSACHUSETTS GENERAL HOSPITAL · PI LEE, PHIL H. · 2020 to 2024
$2.3M
Federated and transfer learning methods for cross-ancestry and cross-phenotype integration of genomic datasetsR01GM148494 · NIGMS · HARVARD UNIVERSITY D/B/A HARVARD SCHOOL OF PUBLIC HEALTH · PI Rui Duan · 2023 to 2026
$1.7M
NIGMS NIH HHS R01 GM148494NIMH NIH HHS R01 MH116037NIMH NIH HHS R01 MH119243NIMH NIH HHS R01 MH124694
6 · The paper itself

Abstract

objectivesWe propose a communication-efficient transfer learning approach (COMMUTE) that effectively incorporates multi-site healthcare data for training a risk prediction model in a target population of interest, accounting for challenges including population heterogeneity and data sharing constraints across sites.

methodsWe first train population-specific source models locally within each site. Using data from a given target population, COMMUTE learns a calibration term for each source model, which adjusts for potential data heterogeneity through flexible distance-based regularizations. In a centralized setting where multi-site data can be directly pooled, all data are combined to train the target model after calibration. When individual-level data are not shareable in some sites, COMMUTE requests only the locally trained models from these sites, with which, COMMUTE generates heterogeneity-adjusted synthetic data for training the target model. We evaluate COMMUTE via extensive simulation studies and an application to multi-site data from the electronic Medical Records and Genomics (eMERGE) Network to predict extreme obesity.

resultsSimulation studies show that COMMUTE outperforms methods without adjusting for population heterogeneity and methods trained in a single population over a broad spectrum of settings. Using eMERGE data, COMMUTE achieves an area under the receiver operating characteristic curve (AUC) around 0.80, which outperforms other benchmark methods with AUC ranging from 0.51 to 0.70.

conclusionCOMMUTE improves the risk prediction in a target population with limited samples and safeguards against negative transfer when some source populations are highly different from the target. In a federated setting, it is highly communication efficient as it only requires each site to share model parameter estimates once, and no iterative communication or higher-order terms are needed.

Indexed as

GenomicsMachine LearningCommunicationComputer SimulationElectronic Health RecordsElectronic health recordsMulti-site studyRisk predictionSynthetic dataTransfer learning

Identifiers

PMID36403757
PMCPMC9868117

What Socratic holds

Textmetadata
LicenceTDM
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.