ArticleJournal of the American Medical Informatics Association : JAMIA2026
SDoH-GPT: using large language models to extract social determinants of health.
Article in Journal of the American Medical Informatics Association : JAMIA, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 10 papers, 1 of them a synthesis that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
10 citing papers in PubMed, 1 synthesis or guideline pooled it.
- Applications of Natural Language Processing and Large Language Models for Social Determinants of Health: Systematic Review.Journal of medical Internet research · 2026Pooled it
- Scalable extraction of social determinants of health from clinical notes in a sepsis cohort using instruction-tuned language models.JAMIA open · 2026Article
- Extracting Social Determinants of Health From Electronic Health Records: Development and Comparison of Rule-Based and Large Language Model Methods.JMIR medical informatics · 2026Article
- Using reasoning LLMs to extract SDOH events from clinical notes.Proceedings. IEEE International Conference on Healthcare Informatics · 2026Article
- ClinNoteAgents: An LLM Multi-Agent System for Predicting and Interpreting Heart Failure 30-Day Readmission from Clinical Notes.AMIA Joint Summits on Translational Science proceedings. AMIA Joint Summits on Translational Science · 2026Article
- Natural language processing-driven knowledge graphs for transformative public health intelligence and research datasets in urgent and emergency care.Frontiers in public health · 2026Article
- Comparing LLM and Fine-Tuned Model Performance on NVDRS Circumstance Extraction with Varying Prompt Complexity.Proceedings. IEEE International Conference on Healthcare Informatics · 2026Article
- Extracting social determinants of health from electronic health records: development and comparison of rule-based and large language models-based methods.medRxiv : the preprint server for health sciences · 2025Article
- A multi-stage large language model framework for extracting suicide-related social determinants of health.Communications medicine · 2025Article
- Unveiling social determinants of health impact on adverse pregnancy outcomes through natural language processing.Scientific reports · 2025Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
14 authors.
Funding
Abstract
objectiveExtracting social determinants of health (SDoHs) from medical notes depends heavily on labor-intensive annotations, which are typically task-specific, hampering reusability and limiting sharing. Here, we introduce SDoH-GPT, a novel framework leveraging few-shot learning large language models (LLMs) to automate the extraction of SDoH from unstructured text, aiming to improve both efficiency and generalizability. MATERIALS AND
methodsSDoH-GPT is a framework including the few-shot learning LLM methods to extract the SDoH from medical notes and the XGBoost classifiers which continue to classify SDoH using the annotations generated by the few-shot learning LLM methods as training datasets. The unique combination of the few-shot learning LLM methods with XGBoost utilizes the strength of LLMs as great few shot learners and the efficiency of XGBoost when the training dataset is sufficient. Therefore, SDoH-GPT can extract SDoH without relying on extensive medical annotations or costly human intervention.
resultsOur approach achieved tenfold and twentyfold reductions in time and cost, respectively, and superior consistency with human annotators measured by Cohen's kappa of up to 0.92. The innovative combination of LLM and XGBoost can ensure high accuracy and computational efficiency while consistently maintaining 0.90+ AUROC scores. DISCUSSION: This study has verified SDoH-GPT on three datasets and highlights the potential of leveraging LLM and XGBoost to revolutionize medical note classification, demonstrating its capability to achieve highly accurate classifications with significantly reduced time and cost.
conclusionThe key contribution of this study is the integration of LLM with XGBoost, which enables cost-effective and high quality annotations of SDoH. This research sets the stage for SDoH can be more accessible, scalable, and impactful in driving future healthcare solutions.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.