Evidence map›Paper›PMID 41580466›Full record

ArticleScientific reports2026

Evaluating the performance of a generative AI model in assessing qualitative health research articles adherence to objective reporting standards.

Aloysius Wei-Yan Chia, Winnie Li-Lian Teo, Yasmin Lynda Munro, Rinkoo Dalan

Abstract read
In one paragraph

Article in Scientific reports, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Aloysius Wei-Yan ChiaDepartment of Endocrinology, Tan Tock Seng Hospital, NHG Health, 11 Jalan Tan Tock Seng, Singapore, 308433, Singapore. aloysius.wy.chia@nhghealth.com.sg.ORCID http://orcid.org/0000-0003-0847-035X
Winnie Li-Lian TeoGroup Education, NHG Health, Annex@National Skin Centre, Level 3, 1 Mandalay Road, Singapore, 308205, Singapore.ORCID http://orcid.org/0000-0003-4592-6344
Yasmin Lynda MunroMedical Library, Lee Kong Chian School of Medicine, Nanyang Technological University, 11 Mandalay Road, Singapore, 308232, Singapore.ORCID http://orcid.org/0000-0002-6261-0037
Rinkoo DalanDepartment of Endocrinology, Tan Tock Seng Hospital, NHG Health, 11 Jalan Tan Tock Seng, Singapore, 308433, Singapore.ORCID http://orcid.org/0000-0001-9769-2696

Funding

Ng Teng Fong Foundation, National Healthcare Group NTF_SRP_P1
6 · The paper itself

Abstract

As qualitative research increasingly informs patient-centred care, rapid assessment of existing evidence to meet research guidelines is needed to inform practice settings. We evaluate the performance of Claude, a generative AI model, in assessing qualitative articles adherence to a consensus-based reporting guideline. The Consolidated Criteria for Reporting Qualitative Research (COREQ), commonly used in qualitative research, is used as a reference criteria list to test the performance of Claude. 15 articles from a systematic scoping review were extracted for analysis. Structured prompts were applied to Claude to evaluate if each criterion in COREQ is met for each article. Two independent reviewers checked model results for concordance and accuracy. The F1, balanced accuracy (BA) scores, Matthews correlation coefficient (MCC) and other performance metrics were tabulated at the criterion, criterion domain, and article level. 4 main categories were identified from performance results, namely: (1) balanced (6/32 criteria, 18.75%), (2) under-reported (2/32, 6.25%), (3) mixed errors (9/32, 28.13%), and (4) information limited (15/32, 46.88%) clusters. Results show heterogeneity amongst different clusters of criteria. While balanced criteria perform consistently across a range of metrics, criteria in under- or over-reported clusters require targeted prompt adjustments. Limited information criteria require a larger sample of articles to verify results. Clearly defined criteria outperformed criteria that were broadly defined or requires interpretation. Segmenting criteria into performance clusters allow researchers to identify areas of incongruence, so that specific strategies to modify prompts may be utilised for any given set of research articles. Customised approaches that are expertly crafted can allow for the rapid extraction of valuable insights that may inform patient-centred recommendations and practice guidelines.

Indexed as

Guideline AdherenceQualitative ResearchGenerative Artificial IntelligenceHumansResearch DesignAdherence to reporting standardsAdherence to standardised checklists and guidelinesAI performance evaluationEvidence synthesisGenerative AIResearch checklists

Identifiers

PMID41580466
PMCPMC12835271

What Socratic holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.