ArticleMedicine2026
Performance of DeepSeek V3 and ChatGPT-4o in answering esophageal cancer-related questions.
Article in Medicine, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
8 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Esophageal cancer remains a significant global health issue. ChatGPT-4o and DeepSeek V3 can provide the public with health-related knowledge about esophageal cancer. This study aimed to evaluate the accuracy of DeepSeek V3 and ChatGPT-4o in responding to health knowledge questions related to esophageal cancer. Fifty-two questions related to esophageal cancer were classified into themes of basic knowledge, diagnosis and molecular biology, management of local and locoregional diseases, management of advanced and metastatic diseases, clinical case analysis and patient frequently asked questions (FAQs). These questions were entered into DeepSeek V3 and ChatGPT-4o to obtain responses, and 2 experienced gastroenterologists independently evaluated the accuracy and temporal stability of each response. Overall, the scores of DeepSeek V3 and ChatGPT-4o on all questions were 4 (3-4), and there was no statistically significant difference between the 2 groups. The final scores of DeepSeek V3 in basic knowledge, diagnosis and molecular biology, management of local and locoregional diseases, management of advanced and metastatic diseases, clinical case analysis, and FAQs were 4 (3-4), 4 (3-4), 4 (3-4), 3 (3-4), 4 (4-4), and 4 (4-4), respectively, while the scores of ChatGPT-4o were 4 (3-4), 3 (2-4), 4 (3-4), 3 (3-4), 4 (4-4), and 4 (4-4), respectively. For temporal stability across 2 independent test runs, DeepSeek V3 presented inconsistent responses on 2 questions, and ChatGPT-4o on 1 question; no statistically significant differences were found in overall and subgroup scores between the 2 runs for both models (all P > .05). ChatGPT-4o and DeepSeek V3 showed favorable accuracy and comprehensive responses to most of the 52 esophageal cancer-related questions in this study, but our findings do not confirm their general reliability for esophageal cancer health information in routine clinical or public use.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.