Evidence mapPaperPMID 41911537Full record

ArticleJournal of medical Internet research2026

Artificial Intelligence Tools for Automating Evidence Synthesis: Scoping Review.

Sashika Harasgama, Helen Pearce, Cameron Appel, Liam Loftus, Helena Painter, Isla Kuhn, Justine Karpusheff, Aji Ceesay, John Ford

Abstract readScoping Review
In one paragraph

Article in Journal of medical Internet research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. Article
  2. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Sashika HarasgamaWolfson Institute of Population Health, Queen Mary University of London, Whitechapel Campus, London, E1 2AD, United Kingdom, +44 (0)20 7882 5555.ORCID http://orcid.org/0009-0003-4099-2470
Helen PearceWolfson Institute of Population Health, Queen Mary University of London, Whitechapel Campus, London, E1 2AD, United Kingdom, +44 (0)20 7882 5555.ORCID http://orcid.org/0009-0006-2625-1432
Cameron AppelWolfson Institute of Population Health, Queen Mary University of London, Whitechapel Campus, London, E1 2AD, United Kingdom, +44 (0)20 7882 5555.ORCID http://orcid.org/0009-0004-7728-5400
Liam LoftusWolfson Institute of Population Health, Queen Mary University of London, Whitechapel Campus, London, E1 2AD, United Kingdom, +44 (0)20 7882 5555.ORCID http://orcid.org/0009-0004-0996-7928
Helena PainterWolfson Institute of Population Health, Queen Mary University of London, Whitechapel Campus, London, E1 2AD, United Kingdom, +44 (0)20 7882 5555.ORCID http://orcid.org/0000-0002-7747-1228
Isla KuhnMedical Library, University of Cambridge, Cambridge, United Kingdom.ORCID http://orcid.org/0000-0002-2879-4020
Justine KarpusheffThe Health Foundation, London, United Kingdom.ORCID http://orcid.org/0000-0003-2797-5898
Aji CeesayThe Health Foundation, London, United Kingdom.ORCID http://orcid.org/0009-0007-5176-3656
John FordWolfson Institute of Population Health, Queen Mary University of London, Whitechapel Campus, London, E1 2AD, United Kingdom, +44 (0)20 7882 5555.ORCID http://orcid.org/0000-0001-8033-7081

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Rapidly and accurately synthesizing large volumes of evidence is a time- and resource-intensive process. Once published, reviews often risk becoming outdated, limiting their usefulness for decision makers. Recent advancements in artificial intelligence (AI) have enabled researchers to automate stages of the evidence synthesis process, from literature searching and screening to data extraction and analysis. As previous reviews on this topic have been published, a significant number of tools have been further developed and evaluated. Furthermore, as generative AI increasingly automates evidence synthesis, understanding how it is studied and applied is crucial, given both its benefits and risks. Objective: This review aimed to map the current landscape of evaluated AI tools used to automate evidence synthesis. Methods: Following the Joanna Briggs Institute methodology for scoping reviews, we searched Ovid MEDLINE, Ovid Embase, Scopus, and Web of Science in February 2025 and conducted a gray literature search in April 2025. We included articles published in any language from January 2021 onward. Two reviewers independently screened citations using Rayyan, and data were extracted based on study design and key AI-related technical features. Results: We identified 7841 unique citations through database searches and 19 records through gray literature searching. A total of 222 articles were included in the review. We identified 65 AI tools and 25 open-source models or machine learning (ML) algorithms that automate parts of or the whole evidence synthesis pathway. A total of 54.1% (n=120) of the studies were published in 2024, reflecting a trend toward researching general-purpose large language models (LLMs) for evidence synthesis automation. The most popular tool studied was generative pretrained transformer models, including its conversational interface ChatGPT (n=70, 31.5%). Moreover, 31.1% (n=69) studied tools automated by traditional ML algorithms. No studies compared traditional ML tools to LLM-based tools. In addition, 61.7% (n=137) and 26.1% (n=58) studied AI-assisted automation of title and abstract screening and data extraction, respectively, the 2 most intensive stages and, therefore, amenable to automation. Technical performance outcomes were the most frequently reported, with only 4.1% (n=9) of studies reporting time- or workload-specific outcomes. Few studies pragmatically evaluated AI tools in real-world evidence synthesis settings. Conclusions: This review comprehensively captures the broad, evolving suite of AI automation tools available to support evidence synthesis, leveraged by increasingly complex AI approaches that range from traditional ML to LLMs. The notable shift toward studying general-purpose generative AI tools reflects how these technologies are actively transforming evidence synthesis practice. The lack of studies in our review comparing different AI approaches for specific automation stages or evaluating their effectiveness pragmatically represents a significant research gap. Optimal tool selection will likely depend on the review topic and methodology and researcher priorities. While they offer potential for reducing workload, ongoing evaluation to mitigate AI bias and to ensure the integrity of reviews is essential for safeguarding evidence-based decision-making.

Indexed as

Artificial IntelligenceGenerative Artificial IntelligenceLarge Language Modelsartificial intelligenceautomationChatGPTevidence synthesislarge language modelsmachine learningsystematic reviews as a topic

Identifiers

PMID41911537
PMCPMC13035263

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.