ArticleScientific reports2026
Evaluation of three artificial intelligence chatbots for generating clinical hematology multiple choice questions for medical students.
Article in Scientific reports, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
3 citing papers in PubMed.
- Generating Image-Based Multiple-Choice Questions with Multimodal Large Language Models: Expert and Psychometric Evaluation.Journal of imaging informatics in medicine · 2026Article
- Evaluating AI Chatbots in Prosthodontics Education: A Quantitative MCQ-Based Assessment.International journal of dentistry · 2026Article
- Quality of Large Language Model-Generated MCQs Across Three Medical Disciplines: An Expert Rater-Based Comparison of Gemini, GPT-4 and Perplexity Pro.Advances in medical education and practice · 2026Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
9 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
The integration of artificial intelligence (AI) into medical education has shown promise in streamlining content creation, yet the reliability and validity of AI-generated assessments remain critical concerns. This study evaluates three AI models-ChatGPT, Perplexity, and DeepSeek-in generating hematology multiple-choice questions (MCQs), focusing on their alignment with clinical guidelines, cognitive complexity, and expert acceptability, to determine their practical utility in medical education. To quantitatively evaluate and compare the performance of three AI models-ChatGPT, Perplexity, and DeepSeek-in generating multiple-choice questions (MCQs) relevant to hematology, with a focus on content validity, cognitive level alignment, and expert acceptance. In this study, each AI model was prompted to generate 50 MCQs across five key hematology topics, following standardized instructions emphasizing guideline alignment and cognitive diversity. Three hematology experts, blinded to question source, independently rated all 150 MCQs on criteria including accuracy, clinical relevance, clarity, distractor plausibility, and overall quality, using a structured rubric. Scores were averaged per model, and questions were categorized by Bloom’s taxonomy level. Acceptance was defined as a total score ≥ 15 out of 25. DeepSeek achieved the highest scores for accuracy (4.7 ± 0.4), clinical relevance (4.8 ± 0.3), and distractor plausibility (4.7 ± 0.4), with a perfect acceptance rate (100%) and no need for revision. Perplexity and ChatGPT also produced clinically relevant questions but required minor revisions (acceptance rates: 96% and 90%, respectively). All models favored higher-order cognitive questions. Knowledge and comprehension questions were limited across all models. AI models, particularly DeepSeek, can efficiently generate high-quality, clinically relevant hematology MCQs suitable for medical education and assessment. While DeepSeek demonstrated superior reliability and required minimal expert revision, all models underrepresented foundational knowledge questions and lacked autonomous image-based item generation. Hybrid human-AI workflows and targeted prompt engineering are recommended to optimize cognitive coverage and ensure educational rigor.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.