Evidence map›Paper›PMID 40966592›Full record

ArticleJournal of medical Internet research2025

Large Language Models' Clinical Decision-Making on When to Perform a Kidney Biopsy: Comparative Study.

Michael Toal, Christopher Hill, Michael Quinn, Ciaran O'Neill, Alexander P Maxwell

Abstract readComparative Study
In one paragraph

Article in Journal of medical Internet research, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 5 papers.

0numbers the graph read from it
0cells of the map it votes in
5citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

5 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
  4. Article
  5. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Michael ToalCentre for Public Health, Royal Victoria Hospital, Queen's University Belfast, Grosvenor Road, Belfast, BT12 6BA, United Kingdom, 44 28 9097 6350.ORCID http://orcid.org/0000-0002-1690-9206
Christopher HillRegional Centre for Nephrology and Transplantation, Belfast City Hospital, Belfast, United Kingdom.ORCID http://orcid.org/0000-0002-1947-8384
Michael QuinnCentre for Public Health, Royal Victoria Hospital, Queen's University Belfast, Grosvenor Road, Belfast, BT12 6BA, United Kingdom, 44 28 9097 6350.ORCID http://orcid.org/0009-0003-2987-924X
Ciaran O'NeillCentre for Public Health, Royal Victoria Hospital, Queen's University Belfast, Grosvenor Road, Belfast, BT12 6BA, United Kingdom, 44 28 9097 6350.ORCID http://orcid.org/0000-0001-7668-3934
Alexander P MaxwellCentre for Public Health, Royal Victoria Hospital, Queen's University Belfast, Grosvenor Road, Belfast, BT12 6BA, United Kingdom, 44 28 9097 6350.ORCID http://orcid.org/0000-0002-6110-7253

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Artificial intelligence (AI) and large language models (LLMs) are increasing in sophistication and are being integrated into many disciplines. The potential for LLMs to augment clinical decision-making is an evolving area of research. Objective: This study compared the responses of over 1000 kidney specialist physicians (nephrologists) with the outputs of commonly used LLMs using a questionnaire determining when a kidney biopsy should be performed. Methods: This research group completed a large online questionnaire for nephrologists to determine when a kidney biopsy should be performed. The questionnaire was co-designed with patient input, refined through multiple iterations, and piloted locally before international dissemination. It was the largest international study in the field and demonstrated variation among human clinicians in biopsy propensity relating to human factors such as sex and age, as well as systemic factors such as country, job seniority, and technical proficiency. The same questions were put to both human doctors and LLMs in an identical order in a single session. Eight commonly used LLMs were interrogated: ChatGPT-3.5, Mistral Hugging Face, Perplexity, Microsoft Copilot, Llama 2, GPT-4, MedLM, and Claude 3. The most common response given by clinicians (human mode) for each question was taken as the baseline for comparison. Questionnaire responses on the indications and contraindications for biopsy generated a score (0-44) reflecting biopsy propensity, in which a higher score was used as a surrogate marker for an increased tolerance of potential associated risks. Results: The ability of LLMs to reproduce human expert consensus varied widely with some models demonstrating a balanced approach to risk in a similar manner to humans, while other models reported outputs at either end of the spectrum for risk tolerance. In terms of agreement with the human mode, ChatGPT-3.5 and GPT-4 (OpenAI) had the highest levels of alignment, agreeing with the human mode on 6 out of 11 questions. The total biopsy propensity score generated from the human mode was 23 out of 44. Both OpenAI models produced similar propensity scores between 22 and 24. However, Llama 2 and MS Copilot also scored within this range but with poorer response alignment to the human consensus at only 2 out of 11 questions. The most risk-averse model in this study was MedLM, with a propensity score of 11, and the least risk-averse model was Claude 3, with a score of 34. Conclusions: The outputs of LLMs demonstrated a modest ability to replicate human clinical decision-making in this study; however, performance varied widely between LLM models. Questions with more uniform human responses produced LLM outputs with higher alignment, whereas questions with lower human consensus showed poorer output alignment. This may limit the practical use of LLMs in real-world clinical practice.

Indexed as

Clinical Decision-MakingKidneyLanguageArtificial IntelligenceBiopsyFemaleHumansLarge Language ModelsMaleSurveys and Questionnairesartificial intelligencechronic kidney diseasedecision supportglomerulonephritishematuriakidney biopsykidney failurelarge language modelsmachine learningnephrologyproteinuriarenal biopsy

Identifiers

PMID40966592
PMCPMC12445783

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.