ReviewJournal of medical Internet research2025
Implementing Large Language Models in Health Care: Clinician-Focused Review With Interactive Guideline.
Review in Journal of medical Internet research, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 18 papers, 2 of them syntheses that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
18 citing papers in PubMed, 2 syntheses or guidelines pooled it.
- Evaluating Large Language Models in Ophthalmology: Systematic Review.Journal of medical Internet research · 2025Pooled it
- Trends in AI-based diagnosis and intervention of metabolic diseases: a bibliometric analysis of the literature from 2000 to 2024.Frontiers in medicine · 2025Pooled it
- Large Language Models for Ankle Fracture Classification and Management Prediction from Routine Clinical Documentation: A Single-Center Exploratory Study.Journal of imaging informatics in medicine · 2026Article
- Performance of large language models as a source of clinical information on bacteriophage therapy.Npj viruses · 2026Article
- Article
- Evaluating the Accuracy of Large Language Models in Selecting Appropriate Statistical Tests for Healthcare Research.Cureus · 2026Article
- Performance comparison of a neuro-symbolic large language model system versus conventional AI models and human experts in cholangitis management.BMC medical informatics and decision making · 2026Article
- IFSO Survey: Use of Large Language Models (LLMs) by Metabolic Bariatric Surgeons and Integrated Health Professionals.Obesity surgery · 2026Article
- Generative AI in perioperative medicine and anesthesiology: ethical integration, educational innovation, and the future of clinical professionalism.Journal of anesthesia · 2026Review
- Program Theory and Core Outcome Set Development for a Technology-Assisted Counseling Intervention in Dementia: Multimethods Study.Journal of medical Internet research · 2026Article
- Large Language Models in Patient Health Communication for Atherosclerotic Cardiovascular Disease: Pilot Cross-Sectional Comparative Analysis.JMIR medical informatics · 2026Article
- A roadmap for medical large language models: a review of foundations, applications, and challenges.Military Medical Research · 2026Review
- Large language models in emergency and critical care medicine: a comprehensive review of applications, challenges, and future directions.Burns & trauma · 2026Review
- Benchmarking the readability, quality, and educational suitability of large language models in communicating pertussis.Frontiers in public health · 2026Article
- Multimodal AI fusion: integrating MRI with PET/CT, histopathology, and liquid biopsy for bone tumor diagnosis.Frontiers in oncology · 2026Review
- Author's Reply: Critical Limitations in Systematic Reviews of Large Language Models in Health Care.Journal of medical Internet research · 2025Article
- Critical Limitations in Systematic Reviews of Large Language Models in Health Care.Journal of medical Internet research · 2025Article
- Prompt Engineering in Clinical Practice: Tutorial for Clinicians.Journal of medical Internet research · 2025Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
3 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
backgroundLarge language models (LLMs) can generate outputs understandable by humans, such as answers to medical questions and radiology reports. With the rapid development of LLMs, clinicians face a growing challenge in determining the most suitable algorithms to support their work.
objectiveWe aimed to provide clinicians and other health care practitioners with systematic guidance in selecting an LLM that is relevant and appropriate to their needs and facilitate the integration process of LLMs in health care.
methodsWe conducted a literature search of full-text publications in English on clinical applications of LLMs published between January 1, 2022, and March 31, 2025, on PubMed, ScienceDirect, Scopus, and IEEE Xplore. We excluded papers from journals below a set citation threshold, as well as papers that did not focus on LLMs, were not research based, or did not involve clinical applications. We also conducted a literature search on arXiv within the same investigated period and included papers on the clinical applications of innovative multimodal LLMs. This led to a total of 270 studies.
resultsWe collected 330 LLMs and recorded their application frequency in clinical tasks and frequency of best performance in their context. On the basis of a 5-stage clinical workflow, we found that stages 2, 3, and 4 are key stages in the clinical workflow, involving numerous clinical subtasks and LLMs. However, the diversity of LLMs that may perform optimally in each context remains limited. GPT-3.5 and GPT-4 were the most versatile models in the 5-stage clinical workflow, applied to 52% (29/56) and 71% (40/56) of the clinical subtasks, respectively, and they performed best in 29% (16/56) and 54% (30/56) of the clinical subtasks, respectively. General-purpose LLMs may not perform well in specialized areas as they often require lightweight prompt engineering methods or fine-tuning techniques based on specific datasets to improve model performance. Most LLMs with multimodal abilities are closed-source models and, therefore, lack of transparency, model customization, and fine-tuning for specific clinical tasks and may also pose challenges regarding data protection and privacy, which are common requirements in clinical settings.
conclusionsIn this review, we found that LLMs may help clinicians in a variety of clinical tasks. However, we did not find evidence of generalist clinical LLMs successfully applicable to a wide range of clinical tasks. Therefore, their clinical deployment remains challenging. On the basis of this review, we propose an interactive online guideline for clinicians to select suitable LLMs by clinical task. With a clinical perspective and free of unnecessary technical jargon, this guideline may be used as a reference to successfully apply LLMs in clinical settings.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.