Evidence map›Paper›PMID 42570960›Full record

ArticleEvidence-based dentistry2026

Can artificial intelligence accurately assess systematic review quality? Benchmarking large language models for AMSTAR 2 appraisal in dental evidence synthesis.

Nirmal Kurian, Joe Mathew Cherian, Ricku Mathew Reji, Kevin George Varghese

Abstract read
PubMed Publisher
In one paragraph

Article in Evidence-based dentistry, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Nirmal KurianChristian Dental College, Ludhiana, Punjab, India. nirmal36@gmail.com.ORCID http://orcid.org/0000-0001-6384-3468
Joe Mathew CherianChristian Dental College, Ludhiana, Punjab, India.ORCID http://orcid.org/0000-0001-8143-4051
Ricku Mathew RejiChristian Dental College, Ludhiana, Punjab, India.
Kevin George VargheseChristian Dental College, Ludhiana, Punjab, India.ORCID http://orcid.org/0000-0002-0120-8159

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

objectiveTo evaluate the accuracy and reliability of three AI platforms ChatGPT, Perplexity, and Google Gemini in assessing the methodological quality of systematic reviews using the AMSTAR 2 checklist, compared with expert manual evaluation in dental research.

methodsA cross-sectional comparative study was conducted to assess the performance of three AI platforms ChatGPT, Perplexity, and Google Gemini in evaluating the methodological quality of 35 systematic reviews using the AMSTAR 2 checklist. Manual assessments by a domain expert served as the reference standard. Each AI system was prompted with a standardized AMSTAR 2 query, and item-level outputs were collected for direct comparison. Key metrics included percentage agreement, error proportions, and inter-rater reliability measured by Cohen's kappa. Error proportions represent the proportion of discordant assessments out of total valid pairwise comparisons across 16 AMSTAR-2 items. Differences between LLM-generated and reference AMSTAR-2 ratings were summarized using effect estimates with corresponding 95% confidence intervals. Comparative performance across platforms was assessed based on confidence-interval overlap rather than hypothesis testing. Results are presented as effect estimates with corresponding 95% confidence intervals, without hypothesis testing or statistical dichotomization. This approach provided a robust and reproducible framework to benchmark AI-assisted quality appraisal in dental evidence synthesis.

resultsAmong 35 systematic reviews assessed, Perplexity demonstrated the highest agreement with expert AMSTAR-2 ratings (error proportion: 19.0%; weighted κ_w: 0.78, 95% CI 0.71-0.85), followed by ChatGPT (error proportion: 22.9%; weighted κ_w: 0.62, 95% CI 0.54-0.70) and Google Gemini (error proportion: 43.9%; weighted κ_w: 0.41, 95% CI 0.33-0.49). Perplexity also achieved the best sensitivity (81.3%, 95% CI 76.5-85.4%) and specificity (82.7%, 95% CI 78.1-86.5%) for correctly identifying high-quality reviews. Non-overlapping 95% confidence intervals suggest meaningful differences in performances among platforms, with Perplexity showing superior agreement across all metrics. Across all platforms, agreement was generally higher for non-critical AMSTAR-2 domains involving clear and structured reporting, whereas performance was weaker for critical domains requiring interpretation of complex methodological details, risk-of-bias considerations, and evidence synthesis procedures.

conclusionsPerplexity demonstrated the highest accuracy and agreement with expert assessments of the methodological quality of systematic reviews, suggesting its potential as a supportive AI tool for AMSTAR-2-based appraisal in dental evidence synthesis. In contrast, systematic biases observed in ChatGPT and Google Gemini underscore the continued need for human oversight to ensure the validity of methodological assessments. Differences in agreement and error proportions were observed across all models when compared with expert AMSTAR-2 evaluations, indicating meaningful variability in methodological appraisal performance, reinforcing that AI-assisted appraisal of systematic review methodology should complement rather than replace expert human judgment in dental research.

Identifiers

What Socratic holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.