Evidence map›Paper›PMID 40380276›Full record

ArticleBMC rheumatology2025

Performance of the Large Language Models in African rheumatology: a diagnostic test accuracy study of ChatGPT-4, Gemini, Copilot, and Claude artificial intelligence.

Yannick Laurent Tchenadoyo Bayala, Wendlassida Joelle Stéphanie Zabsonré/Tiendrebeogo, Dieu-Donné Ouedraogo, Fulgence Kaboré, Charles Sougué, Aristide Relwendé Yameogo, Wendlassida Martin Nacanabo, Ismael Ayouba Tinni, Aboubakar Ouedraogo, Yamyellé Enselme Zongo

Abstract read
In one paragraph

Article in BMC rheumatology, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 8 papers.

0numbers the graph read from it
0cells of the map it votes in
8citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

8 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
  4. Article
  5. Article
  6. Article
  7. Review
  8. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Yannick Laurent Tchenadoyo BayalaDepartment of Rheumatology, Bogodogo University Hospital Center, Sector 51, 14 BP 371, Ouagadougou, Burkina Faso. bayalayannick7991@gmail.com.ORCID http://orcid.org/0009-0004-9095-8948
Wendlassida Joelle Stéphanie Zabsonré/TiendrebeogoDepartment of Rheumatology, Bogodogo University Hospital Center, Sector 51, 14 BP 371, Ouagadougou, Burkina Faso.ORCID http://orcid.org/0000-0001-8098-2697
Dieu-Donné OuedraogoDepartment of Rheumatology, Bogodogo University Hospital Center, Sector 51, 14 BP 371, Ouagadougou, Burkina Faso.ORCID http://orcid.org/0000-0003-2625-2516
Fulgence KaboréDepartment of Rheumatology, Bogodogo University Hospital Center, Sector 51, 14 BP 371, Ouagadougou, Burkina Faso.ORCID http://orcid.org/0000-0001-6541-5352
Charles SouguéDepartment of Internal Medicine, Sourou Sanou University Hospital Center, Bobo-Dioulasso, Burkina Faso.ORCID http://orcid.org/0000-0002-8153-6215
Aristide Relwendé YameogoDepartment of Cardiology, Tengandogo University Hospital Center, Ouagadougou, Burkina Faso.ORCID http://orcid.org/0000-0003-3291-1100
Wendlassida Martin NacanaboDepartment of Cardiology, Bogodogo University Hospital Center, Ouagadougou, Burkina Faso.ORCID http://orcid.org/0009-0002-2047-1830
Ismael Ayouba TinniDepartment of Rheumatology, Bogodogo University Hospital Center, Sector 51, 14 BP 371, Ouagadougou, Burkina Faso.ORCID http://orcid.org/0000-0002-5809-7576
Aboubakar OuedraogoDepartment of Rheumatology, Bogodogo University Hospital Center, Sector 51, 14 BP 371, Ouagadougou, Burkina Faso.ORCID http://orcid.org/0000-0002-7620-5016
Yamyellé Enselme ZongoDepartment of Rheumatology, Bogodogo University Hospital Center, Sector 51, 14 BP 371, Ouagadougou, Burkina Faso.ORCID http://orcid.org/0000-0002-7321-6385

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundArtificial intelligence (AI) tools, particularly Large Language Models (LLMs), are revolutionizing medical practice, including rheumatology. However, their diagnostic capabilities remain underexplored in the African context. To assess the diagnostic accuracy of ChatGPT-4, Gemini, Copilot, and Claude AI in rheumatology within an African population.

methodsThis was a cross-sectional analytical study with retrospective data collection, conducted at the Rheumatology Department of Bogodogo University Hospital Center (Burkina Faso) from January 1 to June 30, 2024. Standardized clinical and paraclinical data from 103 patients were submitted to the four AI models. The diagnoses proposed by the AIs were compared to expert-confirmed diagnoses established by a panel of senior rheumatologists. Diagnostic accuracy, sensitivity, specificity, and predictive values were calculated for each AI model.

resultsAmong the patients enrolled in the study period, infectious diseases constituted the most common diagnostic category, representing 47.57% (n = 49). ChatGPT-4 achieved the highest diagnostic accuracy (86.41%), followed by Claude AI (85.44%), Copilot (75.73%), and Gemini (71.84%). The inter-model agreement was moderate, with Cohen's kappa coefficients ranging from 0.43 to 0.59. ChatGPT-4 and Claude AI demonstrated high sensitivity (> 90%) for most conditions but had lower performance for neoplastic diseases (sensitivity < 67%). Patients under 50 years old had a significantly higher probability of receiving a correct diagnosis with Copilot (OR = 3.36; 95% CI [1.16-9.71]; p = 0.025).

conclusionLLMs, particularly ChatGPT-4 and Claude AI, show high diagnostic capabilities in rheumatology, despite some limitations in specific disease categories. CLINICAL TRIAL NUMBER: Not applicable.

Indexed as

AfricaArtificial intelligenceDiagnostic accuracyLarge Language ModelsRheumatology

Identifiers

PMID40380276
PMCPMC12083132

What Socratic holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.