Evidence map›Paper›PMID 41559279›Full record

ArticleScientific reports2026

Evaluation of three artificial intelligence chatbots for generating clinical hematology multiple choice questions for medical students.

Wiem Boufrikha, Amira Sallem, Baraa Laabidi, Rahma Mallek, Nader Slama, Sabrine Ben Youssef, Ali Majdoub, Nidhal Hadj Salem, Sarra Boukhris

Abstract read
In one paragraph

Article in Scientific reports, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Wiem BoufrikhaHematology Department, Faculty of Medicine of Monastir, University of Monastir, Monastir, Tunisia.
Amira SallemResearch Laboratory "Environment, Inflammation, Signaling and Pathologies" (LR18ES40), Faculty of Medicine of Monastir, University of Monastir, Avicenne Street, 5000, Monastir, Tunisia. amira.sallem2@gmail.com.
Baraa LaabidiHematology Department, University Hospital of Gabes, University of Sfax, Gabes, Tunisia.
Rahma MallekHematology Department, University Hospital of Sfax, University of Sfax, Sfax, Tunisia.
Nader SlamaHematology Department, Faculty of Medicine of Monastir, University of Monastir, Monastir, Tunisia.
Sabrine Ben YoussefDepartment of Paediatric Surgery, Fattouma Bourguiba Hospital, University of Monastir, Monastir, Tunisia.
Ali MajdoubDepartment of Anesthesiology and Perioperative Medicine, Tahar Sfar Hospital, Mahdia, Tunisia.
Nidhal Hadj SalemDepartment of Forensic Medicine, Faculty of Medicine, University of Monastir, Monastir, Tunisia.
Sarra BoukhrisHematology Department, Faculty of Medicine of Monastir, University of Monastir, Monastir, Tunisia.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

The integration of artificial intelligence (AI) into medical education has shown promise in streamlining content creation, yet the reliability and validity of AI-generated assessments remain critical concerns. This study evaluates three AI models-ChatGPT, Perplexity, and DeepSeek-in generating hematology multiple-choice questions (MCQs), focusing on their alignment with clinical guidelines, cognitive complexity, and expert acceptability, to determine their practical utility in medical education. To quantitatively evaluate and compare the performance of three AI models-ChatGPT, Perplexity, and DeepSeek-in generating multiple-choice questions (MCQs) relevant to hematology, with a focus on content validity, cognitive level alignment, and expert acceptance. In this study, each AI model was prompted to generate 50 MCQs across five key hematology topics, following standardized instructions emphasizing guideline alignment and cognitive diversity. Three hematology experts, blinded to question source, independently rated all 150 MCQs on criteria including accuracy, clinical relevance, clarity, distractor plausibility, and overall quality, using a structured rubric. Scores were averaged per model, and questions were categorized by Bloom’s taxonomy level. Acceptance was defined as a total score ≥ 15 out of 25. DeepSeek achieved the highest scores for accuracy (4.7 ± 0.4), clinical relevance (4.8 ± 0.3), and distractor plausibility (4.7 ± 0.4), with a perfect acceptance rate (100%) and no need for revision. Perplexity and ChatGPT also produced clinically relevant questions but required minor revisions (acceptance rates: 96% and 90%, respectively). All models favored higher-order cognitive questions. Knowledge and comprehension questions were limited across all models. AI models, particularly DeepSeek, can efficiently generate high-quality, clinically relevant hematology MCQs suitable for medical education and assessment. While DeepSeek demonstrated superior reliability and required minimal expert revision, all models underrepresented foundational knowledge questions and lacked autonomous image-based item generation. Hybrid human-AI workflows and targeted prompt engineering are recommended to optimize cognitive coverage and ensure educational rigor.

Indexed as

Artificial IntelligenceEducational MeasurementEducation, MedicalHematologyStudents, MedicalGenerative Artificial IntelligenceHumansReproducibility of ResultsArtificial intelligenceChatGPTDeepSeekHematologyMultiple-Choice questionsPerplexity

Identifiers

PMID41559279
PMCPMC12895027

What Socratic holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.