The rubric · 12 criteria
Score every vendor against all of these before a pilot, and score them again after. The red flags are the fastest way to disqualify.
Clinical accuracy & hallucination rate
- Why it matters
- A fabricated finding in a signed note becomes part of the legal record and propagates into coding, billing, and future care decisions.
- How to test
- Run a blinded audit: 100 consecutive encounters, clinician-scored for fabricated findings, omitted findings, and wrong attribution. Demand the vendor's own audit methodology in writing.
- Red flag
- Vendor quotes accuracy as a single percentage with no denominator, or cannot separate omission errors from fabrication errors.
Omission of clinically significant detail
- Why it matters
- Ambient systems fail more often by leaving things out than by making things up — and omissions are far harder to notice on review.
- How to test
- Compare generated notes against full transcripts for a sample of complex, multi-problem encounters. Count clinically significant omissions per note.
- Red flag
- Testing was done only on simple, single-complaint visits.
Edit burden & time-to-sign
- Why it matters
- The entire value proposition is time returned to the clinician. A note requiring heavy editing moves work rather than removing it.
- How to test
- Measure median seconds of edit time and characters changed per note, per specialty, at week 1 and week 12. Track time-to-signature, not just note-generation time.
- Red flag
- Vendor reports satisfaction scores but not measured edit time or time-to-sign.
Specialty and encounter-type coverage
- Why it matters
- Performance varies enormously between a routine primary-care visit, a psychiatric intake, a surgical consult, and an inpatient progress note.
- How to test
- Pilot in your three highest-volume specialties plus your hardest one. Require per-specialty performance data, not an aggregate.
- Red flag
- Only primary-care benchmarks are available.
EHR integration depth
- Why it matters
- A note that must be copy-pasted is a demo, not a workflow. Depth of write-back determines whether the tool disappears into the day.
- How to test
- Verify bidirectional integration: pulls problem list/meds/history for context, writes structured note plus orders and codes back into the chart natively.
- Red flag
- Integration is copy-paste, a browser extension, or a separate app the clinician must switch to.
Coding & billing linkage
- Why it matters
- Documentation is the input to coding. A scribe that produces coding-ready notes captures revenue and audit protection; one that doesn't leaves both on the table.
- How to test
- Compare suggested E/M levels and diagnosis codes against your coders' determinations on a sample. Measure downgrade and denial rates before and after.
- Red flag
- Coding suggestions cannot be traced back to specific supporting text in the note.
Provenance & attribution
- Why it matters
- When a note is challenged — clinically or legally — you must be able to show what was said versus what the model inferred.
- How to test
- Confirm every generated statement can be traced to a transcript segment, and that transcripts are retained per your retention policy.
- Red flag
- No linkage between note text and source audio/transcript.
Privacy, consent & data use
- Why it matters
- Ambient capture records patients and staff continuously. Consent, retention, and whether your data trains the vendor's model are governance decisions, not IT details.
- How to test
- Get in writing: BAA terms, retention windows, de-identification method, whether your data trains shared models, and the patient consent workflow.
- Red flag
- Model training on your data is opt-out rather than opt-in, or buried in terms.
Accent, language & equity performance
- Why it matters
- Speech recognition degrades measurably on accented English, non-native speakers, and some dialects — turning a documentation tool into an equity problem.
- How to test
- Require stratified performance data by language, accent, and patient age. Test with your actual patient mix, including interpreter-mediated visits.
- Red flag
- No stratified data exists, or non-English support is 'on the roadmap'.
Failure mode & downtime behavior
- Why it matters
- Clinicians build workflow dependence quickly. What happens during an outage determines whether that dependence is safe.
- How to test
- Review uptime history and the documented fallback path. Run a tabletop exercise for a four-hour outage during clinic hours.
- Red flag
- No offline capture fallback and no documented degradation plan.
Attestation & oversight model
- Why it matters
- The clinician signing the note holds the liability. The system must make meaningful review possible rather than encourage rubber-stamping.
- How to test
- Check whether uncertain content is flagged for review, whether edits are tracked, and whether the audit trail would satisfy your compliance office.
- Red flag
- Interface nudges toward one-click sign-off with no uncertainty signaling.
Total cost & measurable ROI
- Why it matters
- Per-clinician-per-month pricing is easy to compare; the real cost includes integration, training, audit, and the coding accuracy delta.
- How to test
- Model three-year total cost including integration and audit staff. Tie ROI to measured hours returned, coding capture, and clinician retention.
- Red flag
- ROI case rests entirely on projected burnout reduction with no measurement plan.
The ambient scribe market · 15 vendors
Every ambient clinical documentation vendor of consequence, with the EHR-native options that compete with all of them. Listed alphabetically within tier — no ranking, no paid placement.
Enterprise ambient documentation with deep Epic integration; among the most widely deployed at large academic systems.
The incumbent: DAX Copilot built on Dragon speech recognition, tightly bound to Microsoft and Epic ecosystems.
Ambient documentation plus coding support across a broad specialty range.
Voice assistant and ambient documentation, positioned for both enterprise and independent practices.
Ambient AI assistant with strong multilingual support and rapid clinic-level deployment.
Lightweight scribe aimed at independent clinicians and small practices; fast self-serve adoption.
Ambient documentation with wide international footprint and template flexibility.
Beyond documentation: real-time clinical decision and quality feedback during the encounter.
Ambient documentation with customizable note styles per clinician.
Pioneer of the human-plus-AI scribe model, now shifting toward fully automated notes.
Ambient documentation inside a broader clinical operations and RCM platform.
Scribe and clinical search assistant, notably established in the Canadian market.
Ambient scribe bundled through EHR vendors (eClinicalWorks and others), lowering integration cost.
The EHR vendor's own ambient and generative documentation features — the default competitor to every third-party scribe.
Oracle's ambient documentation agent embedded in Cerner-lineage EHR workflows.
The one thing most evaluations get wrong
Nearly every pilot measures whether clinicians like the tool. Almost none measure omission rate on complex encounters, or edit time at week twelve rather than week one. Satisfaction is real, but it is not evidence of accuracy — and the note is a legal document either way.