Evidence map›Paper›PMID 41963950›Full record

ArticleBMC medical education2026

Evaluating the clinical decision-making performance of large language models in clinically oriented thoracic anatomy scenarios: a comparative evaluation study.

Zeynep Nisa Karakoyun, Mustafa Deniz Yörük, Mehmed Emre Özdemir, Mehmet İlkay Koşar

Abstract readComparative Study
In one paragraph

Article in BMC medical education, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Zeynep Nisa KarakoyunDepartment of Anatomy, Faculty of Medicine, Muğla Sıtkı Koçman University, Muğla, Turkey. zeynepnisakarakoyun@mu.edu.tr.ORCID http://orcid.org/0000-0003-2933-7443
Mustafa Deniz YörükDepartment of Anatomy, Faculty of Medicine, Muğla Sıtkı Koçman University, Muğla, Turkey.ORCID http://orcid.org/0000-0003-3360-9659
Mehmed Emre ÖzdemirDepartment of Artificial Intelligence, Graduate School of Natural and Applied Sciences, Muğla Sıtkı Koçman University, Muğla, Turkey.ORCID http://orcid.org/0009-0003-6230-3576
Mehmet İlkay KoşarDepartment of Anatomy, Faculty of Medicine, Muğla Sıtkı Koçman University, Muğla, Turkey.ORCID http://orcid.org/0000-0002-5773-1838

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundLarge language models are increasingly used as supportive tools in medical education; however, their reliability in anatomically based clinical reasoning remains insufficiently defined. This study aimed to compare the accuracy, depth, and clinical applicability of model-generated responses to thoracic anatomy–based clinical scenarios and to assess inter-rater reliability among expert anatomists.

methodsThis exploratory comparative study evaluated four widely used models—ChatGPT-4o, DeepSeek-V2, Gemini, and Grok—using 20 open-ended thoracic anatomy questions derived from seven clinical scenarios adapted from Moore’s Clinically Oriented Anatomy. Each model generated responses without modification. Two expert anatomists independently assessed all responses using a standardized 10-point scoring system based on anatomical accuracy and clinical relevance. Inter-rater reliability was analyzed using the Intraclass Correlation Coefficient (ICC). Differences in model performance were examined using the Kruskal–Wallis test, with p < 0.05 considered statistically significant.

resultsInter-rater reliability was excellent across all models (ChatGPT-4o: ICC = 1.000; DeepSeek-V2: ICC = 1.000; Gemini: ICC = 0.956; Grok: ICC = 0.977). Significant differences in performance were observed among the models (p < 0.05). Post-hoc analysis demonstrated that Grok achieved the highest median score (7.75), significantly outperforming ChatGPT-4o (4.0), DeepSeek-V2 (4.5), and Gemini (3.0), which showed comparable performance.

conclusionsLarge language models demonstrate variable yet promising potential in supporting clinically oriented thoracic anatomy education. While some models provide more accurate and contextually appropriate responses, expert oversight remains essential. Understanding model-specific strengths and limitations is critical for the safe and responsible integration of AI into biomedical education.

Indexed as

AnatomyClinical Decision-MakingLarge Language ModelsThoraxClinical CompetenceHumansReproducibility of ResultsArtificial intelligenceClinical reasoningLarge language modelsMedical educationThoracic anatomy

Identifiers

PMID41963950
PMCPMC13217646

What Socratic holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.