Evidence map›Paper›PMID 38214966›Full record

ArticleJournal of medical Internet research2024

Automated Paper Screening for Clinical Reviews Using Large Language Models: Data Analysis Study.

Eddie Guo, Mehul Gupta, Jiawen Deng, Ye-Jean Park, Michael Paget, Christopher Naugler

Open access · goldAbstract read
In one paragraph

Article in Journal of medical Internet research, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 98 papers, 8 of them syntheses that pooled it.

0numbers the graph read from it
0cells of the map it votes in
98citing papers in PubMed, 8 pooled it
5.9field-weighted citation impact, top 3% of its field
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

98 citing papers in PubMed, 8 syntheses or guidelines pooled it, 154 citations in OpenAlex.

  1. Pooled it
  2. Pooled it
  3. Pooled it
  4. Pooled it
  5. Pooled it
  6. Pooled it
  7. Pooled it
  8. Pooled it
  9. Article
  10. Review
  11. Review
  12. Article
  13. Article
  14. Article
  15. Article
  16. Article
  17. Article
  18. Article
  19. Article
  20. Article

38 more citing papers are in PubMed but not listed here.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors at 2 institutions in 1 country.

Eddie GuoCumming School of Medicine, University of Calgary, Calgary, AB, Canada.ORCID 0000-0002-7223-0505
Mehul GuptaCumming School of Medicine, University of Calgary, Calgary, AB, Canada.ORCID 0000-0001-7931-0666
Jiawen DengTemerty Faculty of Medicine, University of Toronto, Toronto, AB, Canada.ORCID 0000-0002-8274-6468
Ye-Jean ParkTemerty Faculty of Medicine, University of Toronto, Toronto, AB, Canada.ORCID 0009-0008-1068-8992
Michael PagetCumming School of Medicine, University of Calgary, Calgary, AB, Canada.ORCID 0000-0002-3322-7661
Christopher NauglerCumming School of Medicine, University of Calgary, Calgary, AB, Canada.ORCID 0000-0002-4570-1279
University of Calgary · CAUniversity of Toronto · CA

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundThe systematic review of clinical research papers is a labor-intensive and time-consuming process that often involves the screening of thousands of titles and abstracts. The accuracy and efficiency of this process are critical for the quality of the review and subsequent health care decisions. Traditional methods rely heavily on human reviewers, often requiring a significant investment of time and resources.

objectiveThis study aims to assess the performance of the OpenAI generative pretrained transformer (GPT) and GPT-4 application programming interfaces (APIs) in accurately and efficiently identifying relevant titles and abstracts from real-world clinical review data sets and comparing their performance against ground truth labeling by 2 independent human reviewers.

methodsWe introduce a novel workflow using the Chat GPT and GPT-4 APIs for screening titles and abstracts in clinical reviews. A Python script was created to make calls to the API with the screening criteria in natural language and a corpus of title and abstract data sets filtered by a minimum of 2 human reviewers. We compared the performance of our model against human-reviewed papers across 6 review papers, screening over 24,000 titles and abstracts.

resultsOur results show an accuracy of 0.91, a macro F

conclusionsLarge language models have the potential to streamline the clinical review process, save valuable time and effort for researchers, and contribute to the overall quality of clinical reviews. By prioritizing the workflow and acting as an aid rather than a replacement for researchers and reviewers, models such as GPT-4 can enhance efficiency and lead to more accurate and reliable conclusions in medical research.

Indexed as

Artificial IntelligenceBiomedical ResearchSystematic Reviews as TopicConsensusData AnalysisHumansNatural Language ProcessingProblem SolvingWorkflowabstract screeningChat GPTclassificationextractextractionfree textGPTGPT-4language modellarge language modelsLLMnatural language processingNLPnonopiod analgesiareview methodologyreview methodsscreeningsystematicsystematic reviewunstructured data

Identifiers

PMID38214966
PMCPMC10818236
OpenAlexW4387144848

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.