Observational studyMedicine2024
Evaluating accuracy and reproducibility of ChatGPT responses to patient-based questions in Ophthalmology: An observational study.
Observational study in Medicine, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 15 papers, 1 of them a synthesis that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
15 citing papers in PubMed, 1 synthesis or guideline pooled it.
- Evaluating Large Language Models in Ophthalmology: Systematic Review.Journal of medical Internet research · 2025Pooled it
- The Reliability of Human Evaluation of Large Language Models in Health Care Settings: Scoping Review.Journal of medical Internet research · 2026Article
- Ability of Large Language Models to Answer Patients' Questions and Generate Educational Materials for Uncommon Retinal Conditions.Journal of vitreoretinal diseases · 2026Article
- Large Language Models in Ophthalmology: A Bibliographic Analysis.Turkish journal of ophthalmology · 2026Article
- Comparing the readability of online patient education materials for LASIK and cataract surgery.International ophthalmology · 2026Article
- How Far Have Large Language Models Advanced in Ophthalmology? A Systematic Review of Their Development, Evaluation, and Readiness for Clinical Use.Research square · 2026Article
- Accuracy and Readability of Chat Generative Pre-Trained Transformer-4 Omni in Answering Ophthalmology Patient Questions.Ophthalmology science · 2026Article
- Utility of ChatGPT in generating accurate client handouts for common veterinary internal medicine diseases.Journal of veterinary internal medicine · 2026Article
- Accuracy and Reproducibility of Different Artificial Intelligence Chatbots' Responses to Patient-Based Vitreoretinal Questions: A Comparative Study.Clinical ophthalmology (Auckland, N.Z.) · 2026Article
- Accuracy, readability, and bias of GPT-4o mini responses to oculoplastic patient questions.Frontiers in ophthalmology · 2026Article
- Article
- Comment: Is ChatGPT a Reliable Auxiliary Tool in Basic Life Support Training and Education? A Cross-sectional Study.Indian journal of critical care medicine : peer-reviewed, official publication of Indian Society of Critical Care Medicine · 2025Article
- Is ChatGPT a Reliable Auxiliary Tool in Basic Life Support Training and Education? A Cross-sectional Study.Indian journal of critical care medicine : peer-reviewed, official publication of Indian Society of Critical Care Medicine · 2025Article
- Large language models in the management of chronic ocular diseases: a scoping review.Frontiers in cell and developmental biology · 2025Review
- Evaluating ChatGPT-4 as a digital patient education tool in anesthesia: A multi-rater quality assessment.Digital healthArticle
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
8 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Chat Generative Pre-Trained Transformer (ChatGPT) is an online large language model that appears to be a popular source of health information, as it can provide patients with answers in the form of human-like text, although the accuracy and safety of its responses are not evident. This study aims to evaluate the accuracy and reproducibility of ChatGPT responses to patients-based questions in ophthalmology. We collected 150 questions from the "Ask an ophthalmologist" page of the American Academy of Ophthalmology, which were reviewed and refined by two ophthalmologists for their eligibility. Each question was inputted into ChatGPT twice using the "new chat" option. The grading scale included the following: (1) comprehensive, (2) correct but inadequate, (3) some correct and some incorrect, and (4) completely incorrect. Totally, 117 questions were inputted into ChatGPT, which provided "comprehensive" responses to 70/117 (59.8%) of questions. Concerning reproducibility, it was defined as no difference in grading categories (1 and 2 vs 3 and 4) between the 2 responses for each question. ChatGPT provided reproducible responses to 91.5% of questions. This study shows moderate accuracy and reproducibility of ChatGPT responses to patients' questions in ophthalmology. ChatGPT may be-after more modifications-a supplementary health information source, which should be used as an adjunct, but not a substitute, to medical advice. The reliability of ChatGPT should undergo more investigations.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.