Evidence map›Paper›PMID 42488979›Full record

ArticleJournal of medical Internet research2026

Evaluation Frameworks for Clinical AI Incorporating Validation Strategies, Real-World Applicability, and Ethical Principles: Scoping Review.

Diana Carolina López Medina, Aida Oliveros-Navarro, Nataly Moreno Angel, Andrés Camilo Herrera-Arellano, Marcela Henao-Pérez

Abstract readScoping Review
In one paragraph

Article in Journal of medical Internet research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Diana Carolina López Medina *Research Grupo Infettare, School of Medicine, Universidad Cooperativa Colombia, Calle 50, No 40-74, Bloque A, Medellín, Antioquia, 050012, Colombia, 57 3108406678.ORCID http://orcid.org/0000-0003-2098-7319
Aida Oliveros-Navarro *Research Grupo Infettare, School of Medicine, Universidad Cooperativa Colombia, Calle 50, No 40-74, Bloque A, Medellín, Antioquia, 050012, Colombia, 57 3108406678.ORCID http://orcid.org/0000-0003-4984-7846
Nataly Moreno Angel *Internal Medicine Residency Program, School of Medicine, Universidad Cooperativa de Colombia, Medellín, Colombia.ORCID http://orcid.org/0009-0004-7572-7105
Andrés Camilo Herrera-Arellano *Medicine Program, School of Medicine, Universidad Cooperativa de Colombia, Medellín, Colombia.ORCID http://orcid.org/0009-0009-4718-8497
Marcela Henao-Pérez *Research Grupo Infettare, School of Medicine, Universidad Cooperativa Colombia, Calle 50, No 40-74, Bloque A, Medellín, Antioquia, 050012, Colombia, 57 3108406678.ORCID http://orcid.org/0000-0002-7337-2871

Funding

Bill & Melinda Gates Foundation INV-003688
6 · The paper itself

Abstract

Background: AI shows substantial potential in health care; however, the absence of standardized evaluation frameworks limits its safe and effective clinical implementation because of inconsistent validation requirements and fragmented ethical principles. Existing guidelines vary in structure, methodological rigor, and ethical integration, creating uncertainty. Objective: This study aimed to systematically map, characterize, and critically analyze existing evaluation frameworks for clinical AI, focusing on three core dimensions: methodological rigor, validation strategies (internal validation, including reporting of technical and clinical performance; external validation, including real-world applicability), and alignment with the United Nations Educational, Scientific and Cultural Organization (UNESCO) AI ethical considerations. Methods: A scoping review was conducted following PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines. Six databases (PubMed, Embase, BVS, EBSCOhost, ProQuest, and Sage) and the Enhancing the Quality and Transparency of Health Research Network were searched without language or date restrictions up to February 2026. Eligible documents included peer-reviewed papers, gray literature, and organizational guidelines describing evaluation or reporting frameworks for clinical AI. Editorials, commentaries, and conference abstracts lacking a clearly defined evaluative framework or clinical applicability were excluded. Two reviewers independently screened records and extracted data. Data were extracted across three domains: (1) general characteristics, (2) methodological rigor and validation parameters, and (3) ethical integration and were synthesized using a dot plot-based gap map. Ethical adherence was assessed using a 10-domain UNESCO-based scoring matrix. No formal risk-of-bias assessment was conducted, consistent with scoping review methodology. Results: From 3363 records, 46 frameworks met the inclusion criteria. Mapping revealed a rapidly expanding but fragmented landscape. Most frameworks targeted investigational use (88%), with limited focus on clinical applicability. Frameworks varied in structure, methodology, and scope, with a predominance of reporting guidelines and few validated tools. Most (63%) were developed through multi-institutional collaborations, and 32.6% incorporated transdisciplinary participation. Only 31.8% reported technical metrics (commonly area under the curve, sensitivity, and specificity), and 15.9% provided clinical indicators (eg, predictive values or calibration). Only 11.4% achieved methodological rigor, incorporating validation aligned with intended use, while most relied on partial validation strategies, highlighting a gap between model development and clinical evaluation. Ethical integration was heterogeneous: only 5 frameworks achieved high compliance (≥80%), whereas 4 scored <10%. The most frequently addressed UNESCO principles were awareness and education (71.1%) and transparency and explainability (70%), while human oversight (24.4%) and adaptive governance (33.3%) were least represented. Findings indicate a misalignment between framework design, validation requirements, and clinical implementation. Conclusions: Evaluation frameworks for clinical AI remain heterogeneous and oriented toward investigational contexts. Critical gaps persist in methodological rigor, validation aligned with intended use, and fragmented ethical coverage. These findings highlight the need for standardized, robust, and ethically grounded frameworks to enable safe, reliable, and scalable integration of AI into clinical practice.

Indexed as

Artificial IntelligenceHumansartificial intelligencebioethicsbiomedical technology assessmentclinical decision support systemsevidence-based practicemachine learningscoping reviewsvalidation studies as topic

Identifiers

PMID42488979
PMCPMC13392654

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.