The procurement claim is now testable
For procurement teams looking for AI voice dictation clinical evidence, the most important development is not another demo or another clinician testimonial. It is Lukac et al.’s randomized controlled trial of ambient AI scribes, published in NEJM AI in 2025. The study randomized 238 physicians across roughly 48,000 visits to DAX Copilot, Nabla, or usual care, making it the first randomized trial in a category that has been sold largely on workflow relief and physician well-being promises.[1]
The verdict is mixed in a way that matters. Nabla produced a modest objective reduction in note-writing time. DAX Copilot did not show a statistically significant time reduction. Both products improved physician-reported burnout, task load, and professional fulfillment measures, but those were secondary outcomes rather than the primary endpoint.[1]
| Result | What the RCT found | Procurement implication |
|---|---|---|
| Nabla and documentation time | Note-writing time decreased by 41 seconds, a 9.5% reduction, with a 95% CI from −17.2% to −1.8%.[1] | There is randomized evidence of modest objective time savings for this product in this setting. |
| DAX Copilot and documentation time | No statistically significant reduction in documentation time was found.[1] | The trial does not support treating all ambient scribes as interchangeable time-saving tools. |
| Burnout, task load, and professional fulfillment | Both platforms showed statistically significant improvements on Mini-Z, Physician Task Load, and Professional Fulfillment Index measures.[1] | The well-being signal is more consistent than the productivity signal, but it comes from secondary, subjective endpoints. |

That side-by-side comparison is the part that can get lost in a buying conversation. A vendor can truthfully cite randomized evidence that an ambient scribe improved burnout-related measures. A different sentence can truthfully say one tested product reduced note-writing time. Neither sentence justifies a blanket assumption that ambient AI scribes will produce large, category-wide documentation time savings across a health system.
What the RCT actually measured
The strength of the Lukac trial is its design. Randomization gives this study a different evidentiary status from before-after pilots, satisfaction surveys, and implementation case reports. It also makes the negative and modest findings harder to wave away. If a randomized trial finds a time benefit for one product but not another, the responsible inference is product- and setting-specific, not category-wide.
The primary outcome was time, measured from EHR metadata rather than physician recollection. That distinction is not academic. A physician can feel less depleted at the end of clinic and still spend about the same amount of measurable time in the record. A clinician can also report that a note felt easier to complete even if the total time saved is too small to change staffing models, template policy, or after-hours work expectations.
Nabla’s time effect was real within the trial’s measurement frame but small: 41 seconds less note-writing time, corresponding to a 9.5% reduction, with the confidence interval excluding no effect. DAX Copilot, by contrast, did not produce a statistically significant documentation time reduction in the same trial framework.[1]
The well-being findings were more uniform. Both DAX Copilot and Nabla improved scores on the Mini-Z, Physician Task Load, and Professional Fulfillment Index instruments.[1] Those outcomes should not be dismissed as soft. Burnout is a workforce, quality, access, and retention problem. If an intervention makes clinic feel less punishing for physicians, that may be worth paying for even when the EHR clock does not move much.
But secondary endpoints are not the same thing as a primary trial win. The study gives stronger support to a cautious statement — ambient scribes may reduce perceived burden and improve well-being in outpatient physicians — than to a stronger operational claim that ambient scribes reliably free measurable documentation time. The unblinded nature of the intervention also matters for subjective outcomes: physicians knew whether they were using an AI tool, and expectation effects cannot be fully separated from the experience of actual workflow relief.
The 30% utilization problem is not a footnote

Only about 30% of eligible visits used the assigned AI scribe.[1] For a trial, that is a limitation. For procurement, it is a planning variable.
A budget model that assumes every eligible visit will generate time savings is already out of step with the randomized evidence. If actual use clusters among a subset of physicians, specialties, visit types, or days of the week, the return on investment will follow that clustering. Licenses may be purchased broadly while measurable benefit accrues narrowly.
Low utilization also complicates the interpretation of benefit. The physicians who choose to use an assigned scribe may be more motivated, more comfortable with dictation-style workflows, more burdened at baseline, or better matched to the product’s note structure. The RCT design still protects the main comparison more than an observational rollout would, but adoption behavior remains part of the intervention. Ambient scribe performance is not just model output; it is a coupled system of clinician consent, patient context, room setup, specialty vocabulary, note-review habit, and EHR integration.
This is where many pilots fail to ask the right question. The question is not only whether the software can produce a note. It is whether physicians use it often enough, in the right encounters, with little enough correction work, to change a metric the organization actually cares about.
- For ROI modeling, utilization should be an input, not an afterthought. A pilot should report the share of eligible visits in which the tool was actually used.
- For rollout planning, specialty and visit-type targeting should precede enterprise licensing. A product that helps with conversational outpatient follow-up may not help equally with every clinical workflow.
- For physician selection, baseline documentation burden matters. The highest-benefit users may not be the earliest volunteers, and the earliest volunteers may not represent the median clinician.
- For governance, note-review burden should be measured directly. A shorter drafting phase is less useful if the physician spends the saved time auditing, correcting, or restructuring the output.
Time saved is only meaningful if the work did not move somewhere else
Ambient scribes change the composition of documentation work. The physician may do less initial composing and more review. That can be a good trade if the generated note is accurate, organized, and clinically faithful. It can be a bad trade if the physician has to hunt for subtle errors in a polished paragraph that reads more confidently than it deserves.
The Lukac trial’s time endpoint is therefore important but incomplete as an operational picture. EHR metadata can show whether note-writing time changed. It may not fully capture the mental load of verifying generated content, the interruption pattern of fixing mistakes, or the medicolegal discomfort of signing a note that originated from an opaque summarization process.
That is why the trial’s divergence between objective time and subjective burden is plausible rather than contradictory. A physician may feel less drained because the blank-page problem disappears, because the encounter can feel more conversational, or because the final note is easier to assemble. Those benefits can exist even when the measured time reduction is modest. They just should not be converted automatically into staffing savings, expanded visit capacity, or reduced documentation FTE assumptions.
Accuracy signals deserve more attention than reassurance
The RCT reported clinically significant inaccuracies as “occasional” on 5-point Likert scales for both platforms.[1] That is not a catastrophe signal. It is also not a safety clearance.
The practical issue is that occasional clinically significant inaccuracies can be enough to require a governance response. Ambient scribe errors are not all equivalent. A missing normal review-of-systems phrase is different from a medication error, a negated symptom recorded as present, or a plan detail attached to the wrong condition. The trial result supports the need for physician review; it does not quantify how often the rare, consequential version will appear in broader deployment.
One grade 1, mild adverse event was reported, and no significant patient harm was observed in the trial.[1] That finding should be read within the study’s scale and purpose. A 238-physician trial across roughly 48,000 visits is meaningful for workflow endpoints, but it is not powered to detect rare documentation harms with confidence.[1] Absence of observed major harm is reassuring only to a point.
For an AI governance committee, the minimum response is not to block adoption on the basis of this signal. It is to require a monitoring plan: sampled note audits, error taxonomy, escalation pathways, specialty-specific review, and a clear rule that final responsibility for the note remains with the signing clinician unless policy and law say otherwise. The trial helps identify what to watch. It does not finish the safety assessment.
The observational literature is useful, but it should not outrank the trial
Non-randomized studies are still worth reading because they show what happens when tools enter real delivery systems. They are less reliable for causal inference, especially when physicians opt in, workflows change at the same time, or enthusiasm affects survey response.
Haberle et al.’s Intermountain cohort study of Nuance DAX found no significant productivity gains, and after-hours EHR time increased by 4.69%.[2] That finding is a useful counterweight to the assumption that ambient documentation automatically gives time back. It also shows why organizations should measure after-hours work directly rather than infer it from satisfaction.
Tierney et al.’s Permanente implementation report is valuable for scale. It described more than 2.5 million uses and estimated roughly 18 seconds saved per appointment.[3] That is not nothing at system scale, but it is modest at the encounter level. It is also not the same evidentiary object as a randomized trial.
More favorable estimates from self-selected early-adopter cohorts, including a Mass General Brigham cohort reporting 5.6 minutes saved, should be interpreted in the same frame. Early users often differ from the clinicians who will inherit the tool after enterprise rollout. They may have more compatible workflows, more patience for setup friction, or greater baseline documentation pain. Those cohorts can help identify where an ambient scribe works best, but they are a poor basis for assuming the median physician will benefit to the same degree.
The broader evidence base is still moving. A 2025 systematic review examined AI-powered voice-to-text technology for clinical documentation and quality-of-care outcomes, and a 2026 BMJ Open protocol signals that further synthesis specific to ambient AI scribe documentation is underway.[4][5] That is another reason not to freeze procurement assumptions around one favorable metric from one implementation report.
What a cautious pilot should measure
A defensible pilot does not have to be timid. It does have to be measured in a way that prevents enthusiasm from becoming a spreadsheet artifact.
| Domain | Metric to collect | Why it matters |
|---|---|---|
| Adoption | Percentage of eligible visits using the scribe | The RCT’s roughly 30% use rate means assigned access cannot be treated as actual exposure.[1] |
| Objective time | EHR-derived note-writing time, after-hours EHR time, and time to note closure | These measures separate measurable workflow change from perceived relief. |
| Subjective burden | Burnout, task load, professional fulfillment, and clinician satisfaction | The RCT’s strongest consistent signal was in physician-reported well-being measures.[1] |
| Correction burden | Time spent reviewing and editing generated notes, plus types of edits | The burden may shift from composing to auditing. |
| Accuracy | Sampled note review with clinically significant error categories | The RCT found occasional clinically significant inaccuracies, so local monitoring should be expected.[1] |
| Equity and workflow fit | Use and performance by specialty, visit type, clinician role, and patient communication context | A tool may work well in some outpatient workflows and poorly in others. |
The pilot should also state in advance what decision it is designed to support. A burnout-relief pilot may accept modest or even negligible objective time savings if physician well-being improves enough and safety monitoring is acceptable. A productivity pilot should not rely primarily on survey relief. A capacity-expansion pilot needs a higher bar still, because seconds saved inside note-writing do not automatically convert into additional appointment slots.
The procurement posture supported by the evidence
The Lukac RCT is a milestone because it replaces some of the ambient-scribe fog with randomized evidence. It does not make the category simple. The study supports carefully measured pilots for specific outpatient workflows and physician groups, especially where burnout reduction is an explicit institutional goal. It does not support assuming universal time savings, broad safety assurance, or category-wide ROI.
The most defensible buying position is conditional adoption: choose defined specialties or clinician cohorts, measure utilization and EHR-derived time outcomes, treat well-being as a legitimate but separate endpoint, and audit generated notes for clinically significant errors. If the pilot succeeds, scale toward the workflows that actually used the tool and benefited from it. If it fails, the organization should know whether the failure was adoption, accuracy, integration, visit-type mismatch, or an overoptimistic business case.
Readers tracking regulatory status and incident reports should pair this appraisal with ClinicalMind’s FDA Clearance & Incident Tracker. Terms that often determine how to read studies like this — risk of bias, secondary endpoint, external validity, and adoption versus effectiveness — belong in the ClinicalMind Glossary. This appraisal is meant to support procurement and governance judgment, not clinical guidance for individual patient care.
References
- Ambient AI Scribes in Clinical Practice: A Randomized Trial. NEJM AI. 2025.
- The impact of nuance DAX ambient listening AI documentation: a cohort study. JAMIA / PMC. 2024.
- Ambient AI scribes to alleviate the burden of clinical documentation. NEJM Catalyst. 2023.
- The impact of using AI-powered voice-to-text technology for clinical documentation on quality of care... a systematic review. EClinicalMedicine / ScienceDirect. 2025.
- Use of ambient AI scribe in physicians' clinical documentation: a protocol for a systematic review. BMJ Open. 2026.