Evidence map›Paper›PMID 42414575›Full record

ArticleNPJ digital medicine2026

Benchmarking large language models against practicing clinicians on psychopathological assessment.

Esra Lenz, Joonas Naamanka, Wolfgang Trabert, Ronald Bottlender, Berend Malchow, Andreas Meyer-Lindenberg, Tobias Gradinger, Emanuel Schwarz

Abstract read
In one paragraph

Article in NPJ digital medicine, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Esra LenzHector Institute for Artificial Intelligence in Psychiatry, Central Institute of Mental Health, Medical Faculty Mannheim, Heidelberg University, Mannheim, Germany. esra.lenz@gmail.com.
Joonas NaamankaHector Institute for Artificial Intelligence in Psychiatry, Central Institute of Mental Health, Medical Faculty Mannheim, Heidelberg University, Mannheim, Germany.
Wolfgang TrabertAssociation for Methodology and Documentation in Psychiatry (AMDP), Homburg/Saar, Germany.
Ronald BottlenderAssociation for Methodology and Documentation in Psychiatry (AMDP), Homburg/Saar, Germany.
Berend MalchowAssociation for Methodology and Documentation in Psychiatry (AMDP), Homburg/Saar, Germany.
Andreas Meyer-LindenbergDepartment of Psychiatry and Psychotherapy, Central Institute of Mental Health, Medical Faculty Mannheim, Heidelberg University, Mannheim, Germany.
Tobias Gradinger *Hector Institute for Artificial Intelligence in Psychiatry, Central Institute of Mental Health, Medical Faculty Mannheim, Heidelberg University, Mannheim, Germany.
Emanuel Schwarz *Hector Institute for Artificial Intelligence in Psychiatry, Central Institute of Mental Health, Medical Faculty Mannheim, Heidelberg University, Mannheim, Germany. emanuel.schwarz@zi-mannheim.de.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Psychiatry's reliance on language makes LLMs a natural tool for psychopathological assessment, yet structured, item-level assessments from psychiatric clinical interviews remain under-researched. In this proof-of-concept study, 10 LLMs assessed transcripts of three simulated psychiatric interviews across all 100 items of the Association for Methodology and Documentation in Psychiatry (AMDP) system, benchmarked against 108 early-career clinicians rating full audiovisual recordings, using an expert consensus panel as reference. GPT-5.1 and Gemini-3-Pro-Preview achieved the highest accuracy (0.72; 64th percentile of the clinician distribution) using majority voting across three runs with AMDP definitions as context. GPT-5.1, selected for a marginal advantage, showed per-scenario accuracies of 0.81 (depression), 0.76 (mania), and 0.60 (schizophrenia) versus clinician means of 0.79, 0.68, and 0.58. Clinicians and LLMs showed distinct error profiles: clinicians tended to over-infer symptom presence, whereas LLMs more conservatively flagged items as "not assessable" - most pronounced for observation-dependent items but present even for text-assessable items (19.4% vs. 11.4%, p < 0.001). In post hoc simulated disagreement resolutions (2091 clinician pairs; 35.5% disagreements), LLM and board-certified supervision were associated with more accurate resolutions than unsupervised random clinician selection (p < 0.0002). These proof-of-concept findings require validation in real patient interviews, larger samples, and prospective studies integrating multimodal input.

Identifiers

PMID42414575
PMCPMC13342660

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.