Evidence mapPaperPMID 40705654Full record

ArticleJMIR medical informatics2025

Comparative Analysis of Generative Artificial Intelligence Systems in Solving Clinical Pharmacy Problems: Mixed Methods Study.

Lulu Li, Pengqiang Du, Xiaojing Huang, Hongwei Zhao, Ming Ni, Meng Yan, Aifeng Wang

Abstract readComparative Study
In one paragraph

Article in JMIR medical informatics, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 6 papers, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
6citing papers in PubMed, 1 pooled it
field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

6 citing papers in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
  2. Article
  3. Article
  4. Article
  5. Review
  6. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Lulu LiDepartment of Pharmacy, Fuwai Central China Cardiovascular Hospital, 1 Fuwai Road, Zhengdong New District, Zhengzhou, China, 86 18538298379.ORCID http://orcid.org/0009-0006-4974-132X
Pengqiang DuDepartment of Pharmacy, Fuwai Central China Cardiovascular Hospital, 1 Fuwai Road, Zhengdong New District, Zhengzhou, China, 86 18538298379.ORCID http://orcid.org/0000-0001-6943-485X
Xiaojing HuangDepartment of Pharmacy, Fuwai Central China Cardiovascular Hospital, 1 Fuwai Road, Zhengdong New District, Zhengzhou, China, 86 18538298379.ORCID http://orcid.org/0009-0005-9934-291X
Hongwei ZhaoDepartment of Pharmacy, Fuwai Central China Cardiovascular Hospital, 1 Fuwai Road, Zhengdong New District, Zhengzhou, China, 86 18538298379.ORCID http://orcid.org/0009-0007-7520-7531
Ming NiDepartment of Pharmacy, Fuwai Central China Cardiovascular Hospital, 1 Fuwai Road, Zhengdong New District, Zhengzhou, China, 86 18538298379.ORCID http://orcid.org/0000-0002-9358-8549
Meng YanDepartment of Pharmacy, Fuwai Central China Cardiovascular Hospital, 1 Fuwai Road, Zhengdong New District, Zhengzhou, China, 86 18538298379.ORCID http://orcid.org/0009-0002-8787-7579
Aifeng WangDepartment of Pharmacy, Fuwai Central China Cardiovascular Hospital, 1 Fuwai Road, Zhengdong New District, Zhengzhou, China, 86 18538298379.ORCID http://orcid.org/0009-0004-7948-7073

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Generative artificial intelligence (AI) systems are increasingly deployed in clinical pharmacy; yet, systematic evaluation of their efficacy, limitations, and risks across diverse practice scenarios remains limited. Objective: This study aims to quantitatively evaluate and compare the performance of 8 mainstream generative AI systems across 4 core clinical pharmacy scenarios-medication consultation, medication education, prescription review, and case analysis with pharmaceutical care-using a multidimensional framework. Methods: Forty-eight clinically validated questions were selected via stratified sampling from real-world sources (eg, hospital consultations, clinical case banks, and national pharmacist training databases). Three researchers simultaneously tested 8 different generative AI systems (ERNIE Bot, Doubao, Kimi, Qwen, GPT-4o, Gemini-1.5-Pro, Claude-3.5-Sonnet, and DeepSeek-R1) using standardized prompts within a single day (February 20, 2025). A double-blind scoring design was used, with 6 experienced clinical pharmacists (≥5 years experience) evaluating the AI responses across 6 dimensions: accuracy, rigor, applicability, logical coherence, conciseness, and universality, scored 0-10 per predefined criteria (eg, -3 for inaccuracy and -2 for incomplete rigor). Statistical analysis used one-way ANOVA with Tukey Honestly Significant Difference (HSD) post hoc testing and intraclass correlation coefficients (ICC) for interrater reliability (2-way random model). Qualitative thematic analysis identified recurrent errors and limitations. Results: DeepSeek-R1 (DeepSeek) achieved the highest overall performance (mean composite score: medication consultation 9.4, SD 1.0; case analysis 9.3, SD 1.0), significantly outperforming others in complex tasks (P<.05). Critical limitations were observed across models, including high-risk decision errors-75% omitted critical contraindications (eg, ethambutol in optic neuritis) and a lack of localization-90% erroneously recommended macrolides for drug-resistant Mycoplasma pneumoniae (China's high-resistance setting), while only DeepSeek-R1 aligned with updated American Academy of Pediatrics (AAP) guidelines for pediatric doxycycline. Complex reasoning deficits: only Claude-3.5-Sonnet detected a gender-diagnosis contradiction (prostatic hyperplasia in female); no model identified diazepam's 7-day prescription limit. Interrater consistency was lowest for conciseness in case analysis (ICC=0.70), reflecting evaluator disagreement on complex outputs. ERNIE Bot (Baidu) consistently underperformed (case analysis: 6.8, SD 1.5; P<.001 vs DeepSeek-R1). Conclusions: While generative AI shows promise as a pharmacist assistance tool, significant limitations-including high-risk errors (eg, contraindication omissions), inadequate localization, and complex reasoning gaps-preclude autonomous clinical decision-making. Performance stratification highlights DeepSeek-R1's current advantage, but all systems require optimization in dynamic knowledge updating, complex scenario reasoning, and output interpretability. Future deployment must prioritize human oversight (human-AI co-review), ethical safeguards, and continuous evaluation frameworks.

Indexed as

Artificial IntelligencePharmacy Service, HospitalProblem SolvingGenerative Artificial IntelligenceHumansartificial intelligenceclinical pharmacycomparative analysisDeepSeek-R1generative AI

Identifiers

PMID40705654
PMCPMC12288765

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.