Evidence map›Paper›PMID 41437331›Full record

ArticleBMC medical imaging2025

LLM-powered TNM staging of neuroendocrine tumors from PET/CT reports.

Markus Mergen, Daniel Spitzl, Matthias Eiber, Rickmer F Braren, Lisa Steinhelfer

Abstract read
In one paragraph

Article in BMC medical imaging, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 4 papers.

0numbers the graph read from it
0cells of the map it votes in
4citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

4 citing papers in PubMed.

  1. Article
  2. LLM-powered prostate cancer staging from PSMA-PET/CT reports using PROMISE v2.European journal of nuclear medicine and molecular imaging · 2026
    Article
  3. Article
  4. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Markus Mergen *Institute for Diagnostic and Interventional Radiology, School of Medicine and Health, TUM Klinikum, Klinikum rechts der Isar, Technical University of Munich (TUM), Ismaningerstr. 22, 81675, Munich, Germany.
Daniel Spitzl *Institute for Diagnostic and Interventional Radiology, School of Medicine and Health, TUM Klinikum, Klinikum rechts der Isar, Technical University of Munich (TUM), Ismaningerstr. 22, 81675, Munich, Germany. danieljan.spitzl@mri.tum.de.
Matthias EiberDepartment of Nuclear Medicine, School of Medicine and Health, TUM University Hospital, Technical University Munich, Munich, Germany.
Rickmer F BrarenInstitute for Diagnostic and Interventional Radiology, School of Medicine and Health, TUM Klinikum, Klinikum rechts der Isar, Technical University of Munich (TUM), Ismaningerstr. 22, 81675, Munich, Germany.
Lisa SteinhelferInstitute for Diagnostic and Interventional Radiology, School of Medicine and Health, TUM Klinikum, Klinikum rechts der Isar, Technical University of Munich (TUM), Ismaningerstr. 22, 81675, Munich, Germany.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

purposeImaging reports are essential for the diagnostic evaluation, treatment planning, and follow-up of patients with neuroendocrine tumors (NETs) of the gastroenteropancreatic (GEP) system. The tumor-node metastasis (TNM) classification is a common model for evaluating the prognostic value of tumor patients. However, their traditional free-text format varies in structure, detail, and clarity, leading to inconsistencies and potential omissions of critical information necessary for optimal patient management. Recent advancements in large language models (LLMs) have created new opportunities for automating complex medical assessments, including the extraction of UICC and ENETS staging classifications from imaging reports. This approach aims to improve standardization, enhance clarity, and ensure consistency, ultimately facilitating more effective multidisciplinary clinical decision-making. This study evaluates whether large language models (LLMs) can infer UICC and ENETS TNM stage for GEP‑NETs from PET/CT free‑text reports that contain descriptive findings only (no explicit TNM labels).

methodsWe evaluated several models, including ChatGPT-4o, DeepSeek V3, Claude 3.5 Sonnet, and Gemini 2.0 Flash, on a physician-generated fictitious dataset of 108 PET/CT reports with expert-annotated TNM classifications according to UICC and ENETS criteria. Model performance was assessed through F1-scores, comparing LLM-generated classifications against human expert benchmarks.

resultsAmong the tested models, ChatGPT-4o demonstrated the highest accuracy, achieving microF1 scores of 0.79, 0.99 and 0.99, for T, N and M according to UICC and 0.84, 1.00 and 0.99 respectively, according to ENETS. These results indicate that LLMs have the potential to assist in oncologic staging of NETs, especially offering support for non-specialists in clinical decision-making. However, before integration into routine practice, further prospective validation and rigorous evaluation in real-world settings are necessary.

conclusionThis study underscores the promise of LLMs in oncologic workflows while highlighting the importance of robust benchmarking and clinical validation.

Indexed as

Large Language ModelsNeuroendocrine TumorsPositron Emission Tomography Computed TomographyFemaleHumansMaleNeoplasm StagingClinical decision supportLarge language modelsNeuroendocrine tumorsPET/CTTNM staging

Identifiers

PMID41437331
PMCPMC12838453

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.