Evidence map›Paper›PMID 40644686›Full record

ReviewJournal of medical Internet research2025

Implementing Large Language Models in Health Care: Clinician-Focused Review With Interactive Guideline.

HongYi Li, Jun-Fen Fu, Andre Python

Abstract readReview
In one paragraph

Review in Journal of medical Internet research, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 18 papers, 2 of them syntheses that pooled it.

0numbers the graph read from it
0cells of the map it votes in
18citing papers in PubMed, 2 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

18 citing papers in PubMed, 2 syntheses or guidelines pooled it.

  1. Pooled it
  2. Pooled it
  3. Article
  4. Article
  5. Article
  6. Article
  7. Article
  8. Article
  9. Review
  10. Article
  11. Article
  12. Review
  13. Review
  14. Article
  15. Review
  16. Article
  17. Article
  18. Prompt Engineering in Clinical Practice: Tutorial for Clinicians.Journal of medical Internet research · 2025
    Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

3 authors.

HongYi LiCenter for Data Science, Zhejiang University, Hangzhou, China.ORCID https://orcid.org/0009-0000-0471-8624
Jun-Fen FuSchool of Medicine, Children's Hospital of Zhejiang University, Hangzhou, China.ORCID https://orcid.org/0000-0001-6405-1251
Andre PythonCenter for Data Science, Zhejiang University, Hangzhou, China.ORCID https://orcid.org/0000-0001-8094-7226

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundLarge language models (LLMs) can generate outputs understandable by humans, such as answers to medical questions and radiology reports. With the rapid development of LLMs, clinicians face a growing challenge in determining the most suitable algorithms to support their work.

objectiveWe aimed to provide clinicians and other health care practitioners with systematic guidance in selecting an LLM that is relevant and appropriate to their needs and facilitate the integration process of LLMs in health care.

methodsWe conducted a literature search of full-text publications in English on clinical applications of LLMs published between January 1, 2022, and March 31, 2025, on PubMed, ScienceDirect, Scopus, and IEEE Xplore. We excluded papers from journals below a set citation threshold, as well as papers that did not focus on LLMs, were not research based, or did not involve clinical applications. We also conducted a literature search on arXiv within the same investigated period and included papers on the clinical applications of innovative multimodal LLMs. This led to a total of 270 studies.

resultsWe collected 330 LLMs and recorded their application frequency in clinical tasks and frequency of best performance in their context. On the basis of a 5-stage clinical workflow, we found that stages 2, 3, and 4 are key stages in the clinical workflow, involving numerous clinical subtasks and LLMs. However, the diversity of LLMs that may perform optimally in each context remains limited. GPT-3.5 and GPT-4 were the most versatile models in the 5-stage clinical workflow, applied to 52% (29/56) and 71% (40/56) of the clinical subtasks, respectively, and they performed best in 29% (16/56) and 54% (30/56) of the clinical subtasks, respectively. General-purpose LLMs may not perform well in specialized areas as they often require lightweight prompt engineering methods or fine-tuning techniques based on specific datasets to improve model performance. Most LLMs with multimodal abilities are closed-source models and, therefore, lack of transparency, model customization, and fine-tuning for specific clinical tasks and may also pose challenges regarding data protection and privacy, which are common requirements in clinical settings.

conclusionsIn this review, we found that LLMs may help clinicians in a variety of clinical tasks. However, we did not find evidence of generalist clinical LLMs successfully applicable to a wide range of clinical tasks. Therefore, their clinical deployment remains challenging. On the basis of this review, we propose an interactive online guideline for clinicians to select suitable LLMs by clinical task. With a clinical perspective and free of unnecessary technical jargon, this guideline may be used as a reference to successfully apply LLMs in clinical settings.

Indexed as

Delivery of Health CareLanguageAlgorithmsHumansLarge Language ModelsAIartificial intelligenceclinicaldigital healthlarge language modelLLMLLM review

Identifiers

PMID40644686
PMCPMC12299950

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.