Evidence map›Paper›PMID 39121263›Full record

Observational studyMedicine2024

Evaluating accuracy and reproducibility of ChatGPT responses to patient-based questions in Ophthalmology: An observational study.

Asem A Alqudah, Abdelwahab J Aleshawi, Mohammed Baker, Zaina Alnajjar, Ibrahim Ayasrah, Yaqoot Ta'ani, Mohammad Al Salkhadi, Shaima'a Aljawarneh

Abstract readObservational Study
In one paragraph

Observational study in Medicine, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 15 papers, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
15citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

15 citing papers in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
  2. Article
  3. Article
  4. Article
  5. Article
  6. Article
  7. Article
  8. Article
  9. Article
  10. Article
  11. Article
  12. Comment: Is ChatGPT a Reliable Auxiliary Tool in Basic Life Support Training and Education? A Cross-sectional Study.Indian journal of critical care medicine : peer-reviewed, official publication of Indian Society of Critical Care Medicine · 2025
    Article
  13. Is ChatGPT a Reliable Auxiliary Tool in Basic Life Support Training and Education? A Cross-sectional Study.Indian journal of critical care medicine : peer-reviewed, official publication of Indian Society of Critical Care Medicine · 2025
    Article
  14. Review
  15. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Asem A AlqudahFaculty of Medicine, Jordan University of Science and Technology (JUST), Irbid, Jordan.ORCID 0000-0002-8321-2572
Abdelwahab J AleshawiFaculty of Medicine, Jordan University of Science and Technology (JUST), Irbid, Jordan.
Mohammed BakerFaculty of Medicine, Jordan University of Science and Technology (JUST), Irbid, Jordan.
Zaina AlnajjarFaculty of Medicine, Hashemite University, Zarqa, Jordan.
Ibrahim AyasrahFaculty of Medicine, Jordan University of Science and Technology (JUST), Irbid, Jordan.
Yaqoot Ta'aniFaculty of Medicine, Jordan University of Science and Technology (JUST), Irbid, Jordan.
Mohammad Al SalkhadiFaculty of Medicine, Jordan University of Science and Technology (JUST), Irbid, Jordan.
Shaima'a AljawarnehFaculty of Medicine, Jordan University of Science and Technology (JUST), Irbid, Jordan.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Chat Generative Pre-Trained Transformer (ChatGPT) is an online large language model that appears to be a popular source of health information, as it can provide patients with answers in the form of human-like text, although the accuracy and safety of its responses are not evident. This study aims to evaluate the accuracy and reproducibility of ChatGPT responses to patients-based questions in ophthalmology. We collected 150 questions from the "Ask an ophthalmologist" page of the American Academy of Ophthalmology, which were reviewed and refined by two ophthalmologists for their eligibility. Each question was inputted into ChatGPT twice using the "new chat" option. The grading scale included the following: (1) comprehensive, (2) correct but inadequate, (3) some correct and some incorrect, and (4) completely incorrect. Totally, 117 questions were inputted into ChatGPT, which provided "comprehensive" responses to 70/117 (59.8%) of questions. Concerning reproducibility, it was defined as no difference in grading categories (1 and 2 vs 3 and 4) between the 2 responses for each question. ChatGPT provided reproducible responses to 91.5% of questions. This study shows moderate accuracy and reproducibility of ChatGPT responses to patients' questions in ophthalmology. ChatGPT may be-after more modifications-a supplementary health information source, which should be used as an adjunct, but not a substitute, to medical advice. The reliability of ChatGPT should undergo more investigations.

Indexed as

OphthalmologyHumansInternetReproducibility of ResultsSurveys and Questionnaires

Identifiers

PMID39121263
PMCPMC11315477

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.