ArticleFrontiers in artificial intelligence2025
Assessing the quality of AI-generated clinical notes: validated evaluation of a large language model ambient scribe.
Article in Frontiers in artificial intelligence, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 16 papers, 1 of them a synthesis that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
16 citing papers in PubMed, 1 synthesis or guideline pooled it.
- Evidence on artificial intelligence-assisted clinical documentation and healthcare workers' emotional wellbeing at work: a scoping review.Frontiers in psychology · 2026Pooled it
- An Evaluation of AI-Generated Clinical Notes in the OpenNotes Era: A Thematic Analysis of Clinician Discourse.Journal of evaluation in clinical practice · 2026Article
- Evaluation of an AI Medical Scribe After 236,153 Notes Generated Across Care Levels in a European Health System: Mixed Methods Retrospective Observational Study.JMIR medical informatics · 2026Observational
- Clinicians' rationale for editing ambient AI-drafted clinical notes: persistent challenges and implications for improvement.Journal of the American Medical Informatics Association : JAMIA · 2026Article
- Large Language Model Summarization of Physician-to-Physician Calls for Interhospital Transfer of Patients With ST-Elevation Myocardial Infarction: Observational Study.Journal of medical Internet research · 2026Observational
- From Clinical Encounter to Draft Documentation: A Mechanistic Narrative Review of Ambient Scribe Technology.Cureus · 2026Review
- A Practical Approach to Assessing the Completeness of Electronic Health Records for Medical Research: Data Quality Study.JMIR medical informatics · 2026Article
- Large Language Models for Clinical Narrative Processing: Methods, Applications, and Challenges.Methods and protocols · 2026Article
- Quality of Clinical Notes Created by Ambient Listening Generative AI: Pragmatic Prospective Pilot Study.JMIR medical informatics · 2026Article
- A Bilingual Arabic-English Ambient AI Scribe for Clinical Documentation: Prospective Evaluation Study.JMIR medical informatics · 2026Article
- Large Language Models Using Clinical Text in Pediatrics: A Scoping Review.JAMA network open · 2026Article
- Artificial intelligence chatbots in response to patient's common inquiries about chordoma: A cross-sectional study.Brain & spine · 2026Article
- Real-world evaluation of an ambient AI scribe in Spanish outpatient care after 2.3 million uses: impact on clinician experience, semantic agreement, and workflow efficiency.Frontiers in digital health · 2026Article
- Article
- Breaking barriers: Validation of a Spanish oral health knowledge tool to enhance patient-provider communication.PloS one · 2026Article
- Comparative evaluation of large language models on multiple-choice and image-based rheumatology questions.Rheumatology international · 2025Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
5 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Background: Generative artificial intelligence (AI) tools are increasingly being used as "ambient scribes" to generate drafts for clinical notes from patient encounters. Despite rapid adoption, few studies have systematically evaluated the quality of AI-generated documentation against physician standards using validated frameworks. Objective: This study aimed to compare the quality of large language model (LLM)-generated clinical notes ("Ambient") with physician-authored reference ("Gold") notes across five clinical specialties using the Physician Documentation Quality Instrument (PDQI-9) as a validated framework to assess document quality. Methods: We pooled 97 de-identified audio recordings of outpatient clinical encounters across general medicine, pediatrics, obstetrics/gynecology, orthopedics, and adult cardiology. For each encounter, clinical notes were generated using both LLM-optimized "Ambient" and blinded physician-drafted "Gold" notes, based solely on audio recording and corresponding transcripts. Two blinded specialty reviewers independently evaluated each note using the modified PDQI-9, which includes 11 criteria rated on a Likert-scale, along with binary hallucination detection. Interrater reliability was assessed using within-group interrater agreement coefficient (RWG) statistics. Paired comparisons were performed using Results: Paired analysis of 97 clinical encounters yielded 194 notes (2 per encounter) and 388 paired reviews. Overall, high interrater agreement was observed (RWG > 0.7), with moderate concordance noted in pediatrics and cardiology. Gold notes achieved higher overall quality scores (4.25/5 vs. 4.20/5, Conclusion: LLM-generated Ambient notes demonstrated quality comparable to physician-authored notes across multiple specialties. While Ambient notes were more thorough and better organized, they were also less succinct and more prone to hallucination. The PDQI-9 provides a validated, practical framework for evaluating AI-generated clinical documentation. This quality assessment methodology can inform iterative quality optimization and support the standardization of ambient AI scribes in clinical practice.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.