Evidence map›Paper›PMID 40493530›Full record

ArticleJournal of the American Medical Informatics Association : JAMIA2026

SDoH-GPT: using large language models to extract social determinants of health.

Bernardo Consoli, Haoyang Wang, Xizhi Wu, Song Wang, Xinyu Zhao, Yanshan Wang, Justin Rousseau, Tom Hartvigsen, Li Shen, Huanmei Wu and 4 more

Abstract read
In one paragraph

Article in Journal of the American Medical Informatics Association : JAMIA, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 10 papers, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
10citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

10 citing papers in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
  2. Article
  3. Article
  4. Using reasoning LLMs to extract SDOH events from clinical notes.Proceedings. IEEE International Conference on Healthcare Informatics · 2026
    Article
  5. ClinNoteAgents: An LLM Multi-Agent System for Predicting and Interpreting Heart Failure 30-Day Readmission from Clinical Notes.AMIA Joint Summits on Translational Science proceedings. AMIA Joint Summits on Translational Science · 2026
    Article
  6. Article
  7. Comparing LLM and Fine-Tuned Model Performance on NVDRS Circumstance Extraction with Varying Prompt Complexity.Proceedings. IEEE International Conference on Healthcare Informatics · 2026
    Article
  8. Article
  9. Article
  10. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

14 authors.

Bernardo ConsoliSchool of Information, University of Texas at Austin, Austin, TX 78712, United States.
Haoyang WangSchool of Information, University of Texas at Austin, Austin, TX 78712, United States.
Xizhi WuDepartment of Health Information Management, University of Pittsburgh, Pittsburgh, PA 15261, United States.
Song WangSchool of Information, University of Texas at Austin, Austin, TX 78712, United States.ORCID 0000-0002-8224-0424
Xinyu ZhaoDepartment of Computer Science, University of North Carolina at Chapel Hill, Chapel Hill, NC 27599, United States.
Yanshan WangDepartment of Health Information Management, University of Pittsburgh, Pittsburgh, PA 15261, United States.
Justin RousseauDepartment of Neurology, University of Texas Southwestern Medical Center, Dallas, TX 75390, United States.ORCID 0000-0002-2817-9124
Tom HartvigsenSchool of Data Science, University of Virginia, Charlottesville, VA 22903, United States.
Li ShenDepartment of Biostatistics, Epidemiology and Informatics, University of Pennsylvania, Philadelphia, PA 19104, United States.ORCID 0000-0002-5443-0503
Huanmei WuCollege of Public Health, Temple University, Philadelphia, PA 19122, United States.
Yifan PengDepartment of Population Health Sciences, Weill Cornell Medicine, New York, NY 10065, United States.ORCID 0000-0001-9309-8331
Qi LongDepartment of Biostatistics, Epidemiology and Informatics, University of Pennsylvania, Philadelphia, PA 19104, United States.ORCID 0000-0003-0660-5230
Tianlong ChenDepartment of Computer Science, University of North Carolina at Chapel Hill, Chapel Hill, NC 27599, United States.
Ying DingSchool of Information, University of Texas at Austin, Austin, TX 78712, United States.

Funding

AIM-AHEAD Coordinating Center - All Four CoresOT2OD032581 · OD · UNIVERSITY OF NORTH TEXAS HLTH SCI CTR · PI Paul Avillach, Bettina M. Beech · 2021 to 2026
$168.7M
Closing the loop with an automatic referral population and summarization systemR01LM014306 · NLM · WEILL MEDICAL COLL OF CORNELL UNIV · PI Yifan Peng, Justin Frederick Rousseau · 2023 to 2026
$2.7M
NIH HHS NIH OTA-21-008NIH HHS NIH R01LM014306-01NIH HHS OT2OD032581NIH HHS OTA-21-008NIH HHS R01LM014306-01NLM NIH HHS R01 LM014306
6 · The paper itself

Abstract

objectiveExtracting social determinants of health (SDoHs) from medical notes depends heavily on labor-intensive annotations, which are typically task-specific, hampering reusability and limiting sharing. Here, we introduce SDoH-GPT, a novel framework leveraging few-shot learning large language models (LLMs) to automate the extraction of SDoH from unstructured text, aiming to improve both efficiency and generalizability. MATERIALS AND

methodsSDoH-GPT is a framework including the few-shot learning LLM methods to extract the SDoH from medical notes and the XGBoost classifiers which continue to classify SDoH using the annotations generated by the few-shot learning LLM methods as training datasets. The unique combination of the few-shot learning LLM methods with XGBoost utilizes the strength of LLMs as great few shot learners and the efficiency of XGBoost when the training dataset is sufficient. Therefore, SDoH-GPT can extract SDoH without relying on extensive medical annotations or costly human intervention.

resultsOur approach achieved tenfold and twentyfold reductions in time and cost, respectively, and superior consistency with human annotators measured by Cohen's kappa of up to 0.92. The innovative combination of LLM and XGBoost can ensure high accuracy and computational efficiency while consistently maintaining 0.90+ AUROC scores. DISCUSSION: This study has verified SDoH-GPT on three datasets and highlights the potential of leveraging LLM and XGBoost to revolutionize medical note classification, demonstrating its capability to achieve highly accurate classifications with significantly reduced time and cost.

conclusionThe key contribution of this study is the integration of LLM with XGBoost, which enables cost-effective and high quality annotations of SDoH. This research sets the stage for SDoH can be more accessible, scalable, and impactful in driving future healthcare solutions.

Indexed as

Data MiningElectronic Health RecordsMachine LearningNatural Language ProcessingSocial Determinants of HealthHumansLarge Language Modelsfew-shot learninglarge language modelssocial determinants of healthXGBoost classifier

Identifiers

PMID40493530
PMCPMC12758468

What Socratic holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.