Evidence map›Paper›PMID 41339707›Full record

ArticleScientific reports2025

Comparative performance of ChatGPT-4o, ChatGPT-5, and gemini 2.5 flash on Persian internal medicine subspecialty board exams.

Shahab Sheikhalishahi, Alireza Haddadi, Saina Sadeghipour, Farzad Rafiei, Hamidreza Soltani

Abstract readComparative Study
In one paragraph

Article in Scientific reports, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 6 papers.

0numbers the graph read from it
0cells of the map it votes in
6citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

6 citing papers in PubMed.

  1. Trial
  2. Article
  3. Article
  4. Article
  5. Article
  6. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Shahab SheikhalishahiStudent Research Committee, Shahid Sadoughi University of Medical Sciences, Yazd, Iran.
Alireza HaddadiStudent Research Committee, Shahid Sadoughi University of Medical Sciences, Yazd, Iran.
Saina SadeghipourStudent Research Committee, Shahid Sadoughi University of Medical Sciences, Yazd, Iran.
Farzad RafieiStudent Research Committee, Shahid Sadoughi University of Medical Sciences, Yazd, Iran. farzarafiei@gmail.com.
Hamidreza SoltaniDepartment of Rheumatology, School of Medicine, Shahid Sadoughi University of Medical Sciences, Yazd, Iran. hr.soltan@yahoo.com.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

This study compared the performance of ChatGPT-4o, ChatGPT-5, and Gemini 2.5 Flash on the 2025 Iranian internal medicine subspecialty board examinations. A total of 650 multiple-choice questions from six subspecialties were tested, excluding image-based items. Each question was presented in Persian, and responses were evaluated against the official answer key. Accuracy rates were 68.9% for ChatGPT-4o, 74.5% for ChatGPT-5, and 79.9% for Gemini 2.5 Flash, with Gemini performing significantly better than both ChatGPT versions. ChatGPT-5 also showed a significant improvement over ChatGPT-4o, confirming rapid progress in model development. Subspecialty analysis revealed stronger results in rheumatology and respiratory medicine compared to nephrology, while question type and length had no significant impact on outcomes. An artificial neural network that combined the outputs of all three models reached 81.6% accuracy, slightly exceeding Gemini alone. These findings highlight Gemini-2.5 as the most reliable model for this high-stakes internal medicine exam. The results support the growing role of advanced AI systems as assistants in medical education and clinical practice. However, further research is needed to assess their use in multimodal and real-world clinical tasks.

Indexed as

Educational MeasurementInternal MedicineSpecialty BoardsGenerative Artificial IntelligenceHumansIranNeural Networks, ComputerArtificial intelligence (AI)ChatGPTGeminiInternal medicine

Identifiers

PMID41339707
PMCPMC12796361

What Socratic holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.