Home / AI scribe evaluation

Guide · Updated September 2026

Evaluate the AI scribe like an engineer, not like the demo.

The time savings are real but modest — that word comes from the study authors, not from us. Which means the buying decision isn't really about minutes saved. It's about documentation quality, consent law, and contract terms. Here is what the published evidence actually shows, study by study, and how to run an evaluation a vendor can't steer.

What each study measured — and didn't

Vendors routinely blend these into one claim. They are different studies with different designs.

JAMA, April 2026

~1,800 scribe users vs 6,770 controls across five health systems (Jun 2023–Aug 2025)

EHR time −13 min/day (−3%); documentation time −16 min/day (−10%); +0.49 visits/week

Burnout — the study did not measure it

NEJM AI, December 2025 (RCT)

238 physicians, 14 specialties, randomized head-to-head

One product: −9.5% time-in-note (statistically significant). Another: −1.7%, not significant

Long-term effects; documentation quality

JAMA Network Open (UCSF)

1,565 physicians, ~1.2M encounters

+1.81 RVU/week (+5.8%), ≈ $3,044/yr at 2025 Medicare rates; no rise in denials

Whether gains were more services or better coding — authors could not distinguish. Not 'ROI'

Mass General Brigham

Separate observational study

−21.2% burnout score at 84 days

A different cohort and design than the JAMA studies — don't merge the claims

JMIR Human Factors, July 2025

6 scribe products, standardized audio, Canadian lab study

All six produced omission errors

None erred on names, meds, or numbers — the misses were clinically relevant details

Columns: study · design · what it found · what it did not measure or could not conclude.

The real error mode is omission

In lab testing, every scribe product produced omission errors — details from the visit that never made the note. None fumbled names, medications, or numbers, which is exactly why omissions are dangerous: the note reads clean. A wrong drug name announces itself; a missing symptom doesn't.

From an engineering standpoint this is expected behavior, not a defect you can configure away: the model summarizes audio, and summarization drops content by design. The audio pipeline adds its own losses — crosstalk, accents, room noise. The fix is procedural: run 10+ trial encounters per clinician and have a blinded colleague compare notes against the visit before anyone trusts the output. Repeat after major vendor model updates.

Consent and recording law is catching up

Twelve states clearly require all-party consent to record a conversation — up to thirteen depending on interpretation. A scribe records the visit, so consent belongs in your workflow, not your terms of service.

The litigation has started: a proposed class action filed November 26, 2025 alleges a major scribe vendor and a California medical group auto-inserted consent statements into 100,000+ charts; parallel suits name two other systems. Those are allegations, not findings — but they preview the theory plaintiffs will use. Separately, recordings and transcripts may be part of the designated record set, which makes retention a decision you need to make before rollout, not after.

And the note is still yours: the signing clinician is fully responsible for its contents. Review before signing, amend via addendum, and loop in your malpractice carrier before deployment — they have opinions.

The contract checklist before you sign.

BAA

Signed, covering subcontractors — see our HIPAA guide for what a BAA does and doesn't do.

Retention

What happens to audio and transcripts, and on what schedule. Match it to your record-set decision.

Model training

A written clause on whether your data trains their models. 'We anonymize it' is not a clause.

Exportable transcripts

You can leave with your data. If export is hard, that's your answer.

EHR integration depth

Draft-in-the-chart versus copy-paste changes adoption more than accuracy does.

Patient decline rate

Ask what share of patients refuse recording — one vendor (Commure) reports 7%. Plan the fallback workflow.

Where does this fit in the bigger picture? Scribe selection is one line item in a full readiness assessment — which also tells you whether documentation is even your highest-payoff AI move.

Questions practices ask about scribes

Do patients have to consent to the recording?+

In all-party-consent states, yes — 12 states clearly, up to 13 by interpretation. Even elsewhere, disclosure is good practice and increasingly expected. Build consent into the visit workflow, not the fine print.

Which scribe saves the most time?+

Honestly: the measured differences are small and vendor-specific. In the one head-to-head randomized trial, one product showed a significant reduction in note time and another didn't. Pick on documentation quality, integration, and contract terms — not marketing minutes.

Can the scribe increase billing?+

One large health-system study found about 1.81 more RVUs per week (+5.8%) with no rise in denials — but its authors could not distinguish more services delivered from more complete coding. Treat vendor ROI claims accordingly.

Who is liable for an error in the note?+

The signing clinician, fully. Review before signing, correct via addendum, and talk to your malpractice carrier before rollout.

Is the recording part of the medical record?+

Recordings and transcripts may be part of the designated record set, which makes retention a decision, not an accident. Decide before rollout what you keep and for how long.

Should we pilot before rolling out?+

Yes. Run trial encounters per clinician with blinded review of the notes before trusting the output — omissions are the error mode that won't announce itself.

Deciding on a scribe is easier with the whole map.

The assessment scores documentation against every other AI opportunity in your operation.