Evidence map›Paper›PMID 41028413›Full record

ArticleNPJ digital medicine2025

Optimizing large language models for detecting symptoms of depression/anxiety in chronic diseases patient communications.

Jiyeong Kim, Stephen P Ma, Michael L Chen, Isaac R Galatzer-Levy, John Torous, Peter J van Roessel, Christopher Sharp, Michael A Pfeffer, Carolyn I Rodriguez, Eleni Linos and 1 more

Abstract read
In one paragraph

Article in NPJ digital medicine, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 6 papers, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
6citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

6 citing papers in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
  2. Article
  3. Article
  4. Article
  5. Review
  6. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

11 authors.

Jiyeong KimStanford Center for Digital Health, Department of Medicine, Stanford University, Stanford, CA, USA. jykim3@stanford.edu.
Stephen P MaDivision of Hospital Medicine, Stanford School of Medicine, Stanford, CA, USA.
Michael L ChenStanford Center for Digital Health, Department of Medicine, Stanford University, Stanford, CA, USA.
Isaac R Galatzer-LevyGoogle LLC, Mountain View, CA, USA.
John TorousDivision of Digital Psychiatry, Department of Psychiatry, Beth Israel Deaconess Medical Center, Boston, MA, USA.
Peter J van RoesselDepartment of Psychiatry and Behavioral Sciences, School of Medicine, Stanford University, Palo Alto, CA, USA.
Christopher SharpDivision of Hospital Medicine, Stanford School of Medicine, Stanford, CA, USA.
Michael A PfefferDivision of Hospital Medicine, Stanford School of Medicine, Stanford, CA, USA.
Carolyn I Rodriguez *Department of Psychiatry and Behavioral Sciences, School of Medicine, Stanford University, Palo Alto, CA, USA.
Eleni Linos *Stanford Center for Digital Health, Department of Medicine, Stanford University, Stanford, CA, USA.
Jonathan H Chen *Division of Hospital Medicine, Stanford School of Medicine, Stanford, CA, USA.

Funding

Patient Oriented Research in Vulnerable Populations with Skin DiseaseK24AR075060 · NIAMS · STANFORD UNIVERSITY · PI Eleni Linos · 2019 to 2026
$1.6M
Leveraging Large Language Models and Machine Learning Algorithms to Assess Depression and Anxiety Symptoms and Risks for Patients with Cardiovascular Disease or Diabetes MellitusK01MH137386 · NIMH · STANFORD UNIVERSITY · PI Jiyeong Kim · 2024 to 2026
$406k
NIH HHS K01MH137386NIH HHS K24AR075060
6 · The paper itself

Abstract

Patients with diabetes are at increased risk of comorbid depression or anxiety, complicating their management. This study evaluated the performance of large language models (LLMs) in detecting these symptoms from secure patient messages. We applied multiple approaches, including engineered prompts, systemic persona, temperature adjustments, and zero-shot and few-shot learning, to identify the best-performing model and enhance performance. Three out of five LLMs demonstrated excellent performance (over 90% of F-1 and accuracy), with Llama 3.1 405B achieving 93% in both F-1 and accuracy using a zero-shot approach. While LLMs showed promise in binary classification and handling complex metrics like Patient Health Questionnaire-4, inconsistencies in challenging cases warrant further real-life assessment. The findings highlight the potential of LLMs to assist in timely screening and referrals, providing valuable empirical knowledge for real-world triage systems that could improve mental health care for patients with chronic diseases.

Identifiers

PMID41028413
PMCPMC12485036

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.