ArticleCureus2025
A Comparative Analysis of Artificial Intelligence Platforms: ChatGPT-4o and Google Gemini in Answering Questions About Birth Control Methods.
Article in Cureus, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 4 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
4 citing papers in PubMed.
- Artificial intelligence in obstetrics and gynecology: Evaluating ChatGPT and Google Gemini in answering patient questions.International journal of gynaecology and obstetrics: the official organ of the International Federation of Gynaecology and Obstetrics · 2026Article
- Large Language model (LLM) in temporomandibular disorder education: a comparative study.BMC oral health · 2025Article
- Comparative performance of ChatGPT-4o, ChatGPT-5, and gemini 2.5 flash on Persian internal medicine subspecialty board exams.Scientific reports · 2025Article
- Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
1 author.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Background Birth control methods (BCMs) are often underutilized or misunderstood, especially among young individuals entering their reproductive years. With the growing reliance on artificial intelligence (AI) platforms for health-related information, this study evaluates the performance of ChatGPT-4o and Google Gemini in addressing commonly asked questions about BCMs. Methods Thirty questions, derived from the American College of Obstetrics and Gynecologists (ACOG) website, were posed to both AI platforms. Questions spanned four categories: general contraception, specific contraceptive types, emergency contraception, and other topics. Responses were evaluated using a five-point rubric assessing Relevance, Completeness, and Lack of False Information (RCL). Overall scores were calculated by averaging the rubric scores. Statistical analysis, including the Wilcoxon Signed-Rank test, Friedman test, and Kruskal-Wallis test, was performed to compare metrics. Results ChatGPT-4o and Google Gemini provided high-quality responses to birth control-related queries, with overall scores averaging 4.38 ± 0.58 and 4.37 ± 0.52, respectively, both categorized as "very good" to "excellent." ChatGPT-4o demonstrated higher scores in the lack of false information, based on descriptive statistics (4.70 ± 0.60 vs. 4.47 ± 0.73), while Google Gemini outperformed in relevance, with a statistically significant difference (4.53 ± 0.57 vs. 4.30 ± 0.70, p = 0.035, large effect size). Completeness scores were comparable (p = 0.655). Statistical analyses revealed no significant differences in overall performance (p = 0.548), though Google Gemini demonstrated a potential trend of stronger performance in the "Other Topics" category. Within-model variability showed ChatGPT-4o had more pronounced differences among metrics (moderate effect size, Kendall's W = 0.357), while Google Gemini exhibited smaller variability (Kendall's W = 0.165). These findings suggest that both platforms offer reliable and complementary tools for addressing knowledge gaps in contraception, with nuanced strengths that warrant further exploration. Conclusions ChatGPT-4o and Google Gemini provided reliable and accurate responses to BCM-related queries, with slight differences in strengths. These findings underscore the potential of AI tools, in addressing public health information needs, particularly for young individuals seeking guidance on contraception. Further studies with larger datasets may elucidate nuanced differences between AI platforms.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.