Evidence map›Paper›PMID 41767019›Full record

ArticleCochrane evidence synthesis and methods2026

Using Large Language Models to Address Contextual Questions in Systematic Reviews.

Susanne Hempel, Kimny Sysawang, Haley K Holmer, Erin Tokutomi, Suchitra Iyer, Zhen Wang, Edi Kuhn, Mohammad Hassan Murad

Abstract read
In one paragraph

Article in Cochrane evidence synthesis and methods, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Susanne HempelSouthern California Evidence Review Center University of Southern California Los Angeles California USA.ORCID https://orcid.org/0000-0003-1597-5110
Kimny SysawangSouthern California Evidence Review Center University of Southern California Los Angeles California USA.
Haley K HolmerScientific Resource Center Portland VA Research Foundation Portland Oregon USA.
Erin TokutomiSouthern California Evidence Review Center University of Southern California Los Angeles California USA.
Suchitra IyerEvidence-based Practice Center Program Agency for Healthcare Research and Quality Rockville Maryland USA.
Zhen WangMayo Clinic Evidence-Based Practice Research Program, Mayo Clinic Rochester New York USA.
Edi KuhnScientific Resource Center Portland VA Research Foundation Portland Oregon USA.
Mohammad Hassan MuradMayo Clinic Evidence-Based Practice Research Program, Mayo Clinic Rochester New York USA.ORCID https://orcid.org/0000-0001-5502-5975

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Objectives: Systematic evidence reviews (SERs) produced by the U.S. Agency for Healthcare Research and Quality (AHRQ) Evidence-based Practice Center (EPC) Program use contextual questions to provide context and background information on the topic. There is currently no standardized approach to address contextual questions in systematic reviews. This study explored the use of publicly available large language models (LLMs) in addressing contextual questions. Study Design: Using a set of 20 published and 5 yet to be published SERs, we selected one contextual question per report and used it as a prompt to elicit answers from an LLM (ChatGPT, Bard, Claude, or Perplexity). Two independent reviewers rated the results using a priori established evaluation criteria (https://osf.io/4k3cu/), comparing the response in the SER to LLM-generated responses. The study was guided by six research questions addressing feasibility, validity of content, validity of structure, mistakes, congruence between responses, and incremental validity of using LLMs to address contextual questions. Results: Using minimal prompt engineering produced relevant responses and documented the feasibility of LLM-generated answers to contextual questions. Responses differed in content and format and are not reproducible (e.g., LLMs update regularly), but LLMs were able to produce articulate, clinically plausible, and well-structured responses. We detected few factual errors, contradictions, and no instance of suspected bias, but citations supporting LLM-generated responses could often not be produced or could not be verified ('confabulations'). Congruence with human generated responses varied, with LLM-generated responses providing more background on the topic and SERs providing more nuanced answers in response to the contextual question. Results regarding incremental validity were mixed and may depend on the tool. Conclusion: LLMs are potentially helpful in addressing contextual questions in systematic reviews but human expertise remains essential for using the generated information in a meaningful way.

Indexed as

artificial intelligencecontextlarge language modelssystematic reviews

Identifiers

PMID41767019
PMCPMC12948247

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.