Evidence map›Paper›PMID 42459946›Full record

ArticleJAMIA open2026

Scalable extraction of social determinants of health from clinical notes in a sepsis cohort using instruction-tuned language models.

Diego Salazar, Pankaj Dipankar, Daniel Smolyak, Quynh C Nguyen

Abstract read
In one paragraph

Article in JAMIA open, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Diego SalazarNational Institute of Nursing Research, National Institutes of Health, 9000 Rockville Pike, Bethesda, MD 20814, United States.ORCID https://orcid.org/0000-0001-6013-0008
Pankaj DipankarNational Institute of Nursing Research, National Institutes of Health, 9000 Rockville Pike, Bethesda, MD 20814, United States.ORCID https://orcid.org/0000-0001-9002-4679
Daniel SmolyakNational Institute of Nursing Research, National Institutes of Health, 9000 Rockville Pike, Bethesda, MD 20814, United States.ORCID https://orcid.org/0000-0003-3860-2181
Quynh C NguyenNational Institute of Nursing Research, National Institutes of Health, 9000 Rockville Pike, Bethesda, MD 20814, United States.ORCID https://orcid.org/0000-0003-4745-6681

Funding

Risk and strength: determining the impact of area-level sentiment and protective factors on birth outcomesR01MD015716 · NIMHD · UNIV OF MARYLAND, COLLEGE PARK · PI NGUYEN, THU · 2021 to 2025
$3.6M
Rosie the Chatbot: Leveraging Automated and Personalized Health Information Communication to Reduce Disparities in Maternal and Child HealthR01MD016037 · NIMHD · UNIV OF MARYLAND, COLLEGE PARK · PI NGUYEN, THU, NORELL, ELIZABETH MARIE · 2021 to 2025
$3.3M
HashtagHealthZIANR000043 · NINR · NATIONAL INSTITUTE OF NURSING RESEARCH · PI NGUYEN, QUYNH · 2025 to 2025
$773k
Intramural NIH HHS ZIA NR000043NIMHD NIH HHS R01 MD015716NIMHD NIH HHS R01 MD016037
6 · The paper itself

Abstract

Objectives: Social determinants of health (SDOH) are incompletely captured in structured electronic health records (EHRs) but are frequently documented in unstructured clinical notes. We evaluated large language models (LLMs) for extracting SDOH from clinical text. Materials and Methods: We constructed an adult sepsis cohort from the Medical Information Mart for Intensive Care-IV (Sequential Organ Failure Assessment ≥2) and analyzed clinical notes from 1 year prior to 30 days following suspected infection. Three instruction‑tuned, decoder‑only LLMs (Mistral‑Instruct‑7B-v0.2, DeepSeek‑R1‑Distill‑Qwen‑14B, and GPT‑oss‑20B) were evaluated using structured prompts with predefined label schemas and few‑shot examples. Performance was benchmarked against a clinically validated annotated dataset and compared with a fine‑tuned encoder‑decoder baseline. Macro‑F1 scores were reported. A gold‑standard Intensive Care Unit (ICU) sepsis subset was independently annotated by 3 reviewers to assess domain‑level performance and ensemble strategies. Results: Decoder‑only models outperformed the fine‑tuned encoder‑decoder baseline across SDOH domains. GPT‑oss achieved the highest macro‑F1 score (0.79) compared with Flan‑T5‑XXL (0.57). Prompt refinement substantially improved extraction accuracy. Ensemble majority voting increased robustness across domains, while unanimous agreement yielded high precision but limited coverage. In a subsequent mortality analysis, extracted SDOH did not independently predict 30-day mortality, which was instead associated with established clinical and demographic risk factors. Discussion: Instruction‑tuned decoder‑only LLMs can reliably extract multiclass SDOH from unstructured clinical notes without task‑specific fine‑tuning. Ensemble and agreement‑based strategies provide practical operating points for high‑precision clinical deployment. Conclusion: These findings support the feasibility of leveraging LLMs to enrich EHRs with structured SDOH data, providing a scalable approach for incorporating social context into downstream risk stratification and health outcome prediction.

Indexed as

clinical noteslarge language modelssepsis mortalitysocial determinants of health (SDOH)

Identifiers

PMID42459946
PMCPMC13371764

What Socratic holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.