Evidence mapPaperPMID 40460423Full record

ArticleJournal of medical Internet research2025

Predicting 30-Day Postoperative Mortality and American Society of Anesthesiologists Physical Status Using Retrieval-Augmented Large Language Models: Development and Validation Study.

Ying-Hao Chen, Shanq-Jang Ruan, Pei-Fu Chen

Registry-linked trialAbstract readValidation Study
In one paragraph

Article in Journal of medical Internet research, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. It is linked to trial NCT07696221 (Comparison of Clinical Assessment and Large Language Models in Preoperative Risk Classification), which is not on this map. Cited by 4 papers.

0numbers the graph read from it
0cells of the map it votes in
4citing papers in PubMed
field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

NCT07696221 not yet recruitingnot on this mapstarted 2026, after this paper: background citation

Comparison of Clinical Assessment and Large Language Models in Preoperative Risk Classification: A Retrospective Analysis of ChatGPT, DeepSeek, Gemini, and Claude in ASA Physical Status Classification

TypeobservationalSponsorMarmara University Pendik Training and Research HospitalRan2026 to 2026Enrolled350ConditionsAnesthesia, Preoperative Risk Prediction, Preoperative Risk Assessment
3 · Its place in the literature

Who cites it

4 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
  4. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

3 authors.

Ying-Hao ChenDepartment of Electronic and Computer Engineering, National Taiwan University of Science and Technology, Taipei, Taiwan.ORCID https://orcid.org/0009-0008-3761-6198
Shanq-Jang RuanDepartment of Electronic and Computer Engineering, National Taiwan University of Science and Technology, Taipei, Taiwan.ORCID https://orcid.org/0000-0003-1075-8512
Pei-Fu ChenDepartment of Anesthesiology, Far Eastern Memorial Hospital, New Taipei City, Taiwan.ORCID https://orcid.org/0000-0002-0192-5377

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundAccurately assessing perioperative risk is critical for informed surgical planning and patient safety. However, current prediction models often rely on structured data and overlook the nuanced clinical reasoning embedded in free-text preoperative notes. Recent advances in large language models (LLMs) have opened opportunities for harnessing unstructured clinical data, yet their application in perioperative prediction remains limited by concerns about factual accuracy. Retrieval-augmented generation (RAG) offers a promising solution-enhancing LLM performance by grounding outputs in domain-specific knowledge sources, potentially improving both predictive accuracy and clinical interpretability.

objectiveThis study aimed to investigate whether integrating LLMs with RAG can improve the prediction of 30-day postoperative mortality and American Society of Anesthesiologists (ASA) physical status classification using unstructured preoperative clinical notes.

methodsWe conducted a retrospective cohort study using 24,491 medical records from a tertiary medical center, including preoperative anesthesia assessments, discharge summaries, and surgical information. To extract clinical insights from free-text data, we used the LLaMA 3.1-8B language model with RAG, using MedEmbed for text embedding and Miller's Anesthesia as the primary retrieval source. We evaluated model performance under various configurations, including embedding models, chunk sizes, and few-shot prompting. Machine learning (ML) models, including random forest, support vector machines (SVM), Extreme Gradient Boosting (XGBoost), and logistic regression, were trained on structured features as baselines.

resultsA total of 520 (2.1%) patients experienced in-hospital 30-day postoperative mortality. The ASA physical status distribution was as follows: class I: 535 (2.2%); class II: 15,272 (62.4%); class III: 8024 (32.8%); class IV: 606 (2.5%); and class V: 54 (0.22%). For 30-day postoperative mortality prediction, the LLaMA‑RAG model achieved an F

conclusionsThe LLaMA-RAG model significantly improved the prediction of postoperative mortality and ASA classification, especially for rare high-risk cases. By grounding outputs in domain knowledge, retrieval-augmented generation enhanced both accuracy and prompt‑driven interpretability over ML and ablation models-highlighting its promise for real-world clinical decision support.

Indexed as

Postoperative ComplicationsAgedAnesthesiologistsFemaleHumansLanguageLarge Language ModelsMaleMiddle AgedPostoperative PeriodRetrospective StudiesSocieties, MedicalUnited Statesfew-shot promptingmachine learningperioperative careprediction modelsurgical risk stratificationunstructured clinical data

Identifiers

PMID40460423
PMCPMC12174870

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.