TOPLINE:
AI-generated clinical notes received lower quality scores than human-produced notes across five standardized cases in primary care, with the largest deficits noted in being thorough, organized, and useful.
METHODOLOGY:
- Researchers from the Veterans Health Administration (VHA) in the US conducted a cross-sectional evaluation comparing the quality of AI-generated and human-produced clinical notes from standardized primary care cases.
- Five audio-recorded standardized patient scenarios were created to represent common clinical encounters, with audio played at a fixed volume through standardized speakers to replicate real-world ambient conditions. These audio recordings incorporated challenges, such as background noise, nonnative accents, and whether participants were wearing surgical masks.
- The five scenarios in primary care were a new patient establishing care with newly diagnosed diabetes, the evaluation of acute low back pain, the assessment of chest pain, a pharmacy consultation, and a follow-up with a nurse care manager for heart failure.
- Each of the 11 AI vendors generated clinical notes in subjective, objective, assessment, and plan formats from the five audio files, and three human clinicians per case (including physicians, pharmacists, and nurse care managers, depending on the scenario) listened to the recordings once and independently produced comparison notes.
- Blinded raters assessed all notes using the nine-item modified Physician Documentation Quality Instrument (PDQI-9), which measures 10 domains of note quality: accuracy, thoroughness, usefulness, organization, comprehensiveness, succinctness, synthesis, internal consistency, and freedom from hallucination and bias.
TAKEAWAY:
- Across all five clinical cases involving 55 AI-generated notes (330 ratings) and 15 human-produced notes (87 ratings), human-produced notes received higher overall modified PDQI-9 scores than AI-generated notes, with the magnitude of difference varying by case.
- The largest difference in quality was observed in the case of acute low back pain with substantial background noise, in which the mean predicted score was 43.8 for the human-produced notes compared with 20.3 for the AI-generated notes, representing a difference of -23.5 (P ≤ .001).
- In pooled analysis, AI-generated notes scored lower across all 10 domains, with notable deficits seen in being thorough (-1.23; 95% CI, -1.82 to -0.65), organized (-1.06; 95% CI, -1.65 to -0.47), and useful (-1.03; 95% CI, -1.61 to -0.44).
- Smaller differences between AI-generated and human-produced notes were observed in domains specific to AI-related concerns of bias (-0.7; 95% CI, -1.32 to -0.14) and hallucination (-0.87; 95% CI, -1.46 to -0.29).
IN PRACTICE:
“For clinicians, AI scribes should be regarded as tools for generating draft documentation that requires review and editing, rather than as a substitute for clinician-authored notes,” the authors of the study wrote.
“As ambient AI scribes continue to reshape the clinical documentation landscape, it is essential to use evaluative frameworks that accurately reflect real-world performance, promote the development of documentation policies that prioritize patient care over billing requirements, and systematically incorporate patient perspectives into assessments of quality,” experts wrote in an accompanying editorial.
SOURCE:
The study was led by Ashok Reddy, MD, MSc, of the Center of Innovation for Veteran-Centered and Value-Driven Care in the VA Puget Sound Health Care System in Seattle. It was published online on April 17, 2026, in the Annals of Internal Medicine. The editorial was written by Aaron A. Tierney, PhD, and Kristine Lee, MD.
LIMITATIONS:
The evaluation used a limited number of standardized primary care cases, which may not capture all the complexities encountered in real-world practice. The encounters were simulated and not actual patient visits. Notes were produced by clinicians outside traditional clinical experience and may not represent the quality of notes produced when under pressure to finish documentation, and the notes were not generated by professional scribes.
DISCLOSURES:
The Primary Care Analytics Team through the VHA Office of Primary Care provided support for this study. Two authors disclosed having employment with the US Department of Veterans Affairs.
This article was created using several editorial tools, including AI, as part of the process. Human editors reviewed this content before publication.
Admin_Adham