ArticleRadiological physics and technology2026
Enhancing pancreatic cancer staging with large language models: the role of retrieval-augmented generation.
Article in Radiological physics and technology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
2 citing papers in PubMed.
- Evaluation Methods for Inference-Time Retrieval-Augmented and Graph Retrieval-Augmented Large Language Models in Health Care: Scoping Review.Journal of medical Internet research · 2026Article
- Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
12 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Retrieval-augmented generation (RAG) is an emerging technique that enhances large language models (LLMs) by retrieving relevant information from reliable external knowledge (REK). Despite its success in various language processing tasks, the clinical application of RAG in radiology is still novel. To evaluate the utility of RAG in radiological staging tasks, we compared the performance of NotebookLM, a RAG-equipped LLM (RAG-LLM), with its internal model, Gemini 2.0 Flash. Using Japan’s current pancreatic cancer staging guideline as REK, we compared three LLM settings—(1) NotebookLM with REK (REK+/RAG+), (2) Gemini 2.0 Flash with REK (REK+/RAG−), and (3) Gemini 2.0 Flash without REK (REK−/RAG−)—in staging 100 fictional pancreatic cancer cases based on CT findings. Staging tasks included assessment of TNM classification, local invasion factors, and resectability classification. The REK+/RAG+ group achieved 70% staging accuracy, outperforming REK+/RAG− (38%) and REK−/RAG− (35%). For TNM classification, REK+/RAG+ reached 80% accuracy, compared to 55% and 50% in REK+/RAG − and REK−/RAG−, respectively. NotebookLM also presented retrieved REK excerpts as rationale, with a retrieval accuracy of 92%. These results suggest that the RAG system implemented in NotebookLM improves LLM staging performance by enabling access to up-to-date medical guidelines. Furthermore, its ability to retrieve and present source evidence enhances transparency and allows users to verify the reliability of model outputs. This study highlights the potential of a specific RAG system (NotebookLM) to support pancreatic cancer staging by combining clinical language interpretation with direct reference to authoritative guidelines, in an idealized and controlled proof-of-concept experimental setting.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.