Evidence map›Paper›PMID 42747656›Full record

ArticleNeurosurgical review2026

Inter-Model Variability of ChatGPT, Gemini and Claude in Treatment Recommendations for Unruptured Intracranial Aneurysms.

Anish Narayan, Frederick Mariajoseph, Malik Farooq, Adrian Praeger, Ronil Chandra, Idrees Sher, Lee-Anne Slater, Calvin Gan, Andrew Gauden, Hamed Asadi and 1 more

Abstract read
In one paragraph

Article in Neurosurgical review, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

11 authors.

Anish NarayanFaculty of Medicine, Monash University, Clayton, Australia. anar0024@student.monash.edu.ORCID https://orcid.org/0009-0009-5326-6949
Frederick MariajosephDepartment of Neurosurgery, Monash Health, Clayton, Australia.
Malik FarooqFaculty of Medicine, Monash University, Clayton, Australia.
Adrian PraegerDepartment of Neurosurgery, Monash Health, Clayton, Australia.
Ronil ChandraMonash Health Imaging, Monash Health, Melbourne, Australia.
Idrees SherDepartment of Neurosurgery, Monash Health, Clayton, Australia.
Lee-Anne SlaterMonash Health Imaging, Monash Health, Melbourne, Australia.
Calvin GanMonash Health Imaging, Monash Health, Melbourne, Australia.
Andrew GaudenDepartment of Neurosurgery, Monash Health, Clayton, Australia.
Hamed AsadiDepartment of Radiology, Austin Health, Heidelberg, VIC, 3084, Australia.
Justin MooreDepartment of Neurosurgery, Monash Health, Clayton, Australia.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Patients increasingly consult large language models (LLMs) before specialist review, yet whether frontier models agree with one another in nuanced clinical domains such as unruptured intracranial aneurysm (UIA) management remains uncharacterised. We therefore quantified inter-model variability across ChatGPT, Gemini and Claude, anchored against neurovascular multidisciplinary team (MDT) consensus and the Unruptured Intracranial Aneurysm Treatment Score (UIATS). Sixty-seven UIA cases referred to our neurovascular service (January-December 2025) were retrospectively analysed. De-identified clinical vignettes were submitted to Claude Opus 4.6, ChatGPT-5.4 and Gemini 3 Pro Thinking, each run five times. Within-model reproducibility was assessed using Fleiss' κ; inter-model agreement via Cohen's κ and McNemar's test; and each LLM's majority-vote anchored against MDT and UIATS using the same methods. Within-model reproducibility was almost perfect (Fleiss' κ 0.837-0.860). Pairwise inter-model agreement was asymmetric: ChatGPT-Gemini behaved near-identically (Cohen's κ = 0.850, 95% CI 0.71-0.97), whereas Claude diverged from both (κ = 0.688 and 0.667). Recommendations were non-unanimous in 13/67 cases (19.4%); Claude was the sole outlier in 8/13 (conservative in 7). Gemini was significantly more pro-treatment than Claude (McNemar p = 0.0117). Against MDT, Gemini and ChatGPT showed significant pro-treatment propensity (p = 0.0022, p = 0.0153). Claude was significantly more conservative than UIATS (p = 0.0162). Frontier LLMs are highly reproducible internally but diverge from one another asymmetrically. ChatGPT and Gemini behave near-identically while Claude diverges conservatively. Clinicians should anticipate AI-driven treatment expectations and future work should explore prompting, patient sentiment, and multicentre moderators.

Indexed as

Intracranial AneurysmFemaleGenerative Artificial IntelligenceHumansLarge Language ModelsMaleMiddle AgedReproducibility of ResultsRetrospective StudiesArtificial intelligenceInter-model variabilityLarge language modelsMultidisciplinary teamUnruptured intracranial aneurysms

Identifiers

PMID42747656
PMCPMC13582248

What Socratic holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.