20 of 20. That's how many AI scribe vendors approved for use across Ontario's public health system produced at least one factual error, fabrication, or missing clinical detail when the province's own evaluators tested them before approval (Office of the Auditor General of Ontario, "Use of Artificial Intelligence in the Ontario Government," Special Report, May 12, 2026). Every vendor that made it onto the approved list had a documented accuracy problem.

This is a Canadian government audit, not a US one, and no practice in Manchester, Ohio buys software through Supply Ontario. But it's the first systematic, comparative accuracy test of commercial AI scribe vendors that most independent practice owners will ever see, because nothing like it has been run in the US. The pattern it exposes, that regulatory approval and clinical accuracy are two different things, applies just as directly to a physical therapy practice evaluating a scribe off a vendor's sales deck.

How the test actually worked

Supply Ontario ran a formal request-for-bid process to approve AI scribe vendors for use across the province's health-care system, in consultation with OntarioMD, Ontario Health, and the Ministry of Health. As part of that process, all 20 bidding vendors were given the same two simulated recordings of a conversation between a health-care professional and a patient. Each vendor's system transcribed the conversation and generated a structured clinical note. Medical professionals from OntarioMD and Ontario Health then reviewed every note for accuracy and completeness.

So what for you: this was not a stress test designed to catch vendors out. It was two standard, simulated conversations, run once, under conditions vendors knew were part of a competitive evaluation. Real-world accuracy, on a full patient day with background noise, interruptions, and genuine clinical complexity, has no obvious reason to be better than what a controlled procurement test found.

The three ways every vendor failed

The Auditor General's evaluators grouped the errors into three categories, and logged which vendors hit each one.

Hallucinations. Nine of 20 vendors (45%) fabricated information outright, including suggestions for a patient's treatment plan, such as a referral to therapy or an order for blood tests, that were never mentioned in the recording. Some of the fabricated content could have changed clinical decisions directly: notes claimed "no masses found" or described anxiety symptoms that were never discussed in the conversation at all.

Incorrect information. Twelve of 20 vendors (60%) generated notes that recorded a different drug than the one the doctor actually prescribed during the simulated conversation. Not a spelling variant. A different medication.

Missing or incomplete information. Seventeen of 20 vendors (85%) missed key details about the patient's mental health in at least one of the two tests, despite those details being clearly stated in the recording. Six of 20 (30%) missed them across both tests, fully or partially.

So what for you: a wrong drug on a chart and a missed mental health detail are not equally serious in every case, but both are exactly the kind of error a rushed clinician, trusting the AI output because it reads fluently, is least likely to catch on a quick read-through.

Both error types map directly onto what a physical therapy note actually contains. Every physio intake records current medications, because drug interactions and side effects change how a treatment plan gets built. And every physio evaluation worth the name screens for psychosocial "yellow flags," the mental health and fear-avoidance factors that predict whether a patient recovers on schedule or stalls. A scribe that gets the drug wrong or drops the mental health detail isn't making a cosmetic error in a physio chart. It's corrupting the two categories of information a treatment plan depends on most.

Why 100% failure still got every vendor approved

Here's the part that should change how a practice owner reads any vendor's "compliant" or "approved" badge. In Supply Ontario's scoring model, "accuracy of medical notes generated" accounted for just 4% of a vendor's total possible score, 20 points out of 530. "Domestic presence in Ontario," a criterion with nothing to do with clinical safety, was weighted highest of all at 30%, or 159 points. Security controls and bias testing scored even lower than accuracy.

More striking: no minimum passing score applied to the accuracy, security, or bias criteria at all. A vendor could score zero on accuracy, zero on security, and zero on bias controls, and still be approved as a supplier by hitting the overall aggregate threshold of 371 out of 530 points. The Auditor General's report recommends Supply Ontario fix exactly this in future procurements.

So what for you: "government-tested" and "government-approved" sound like reassurance. This report shows they can both be true of a product that failed a basic accuracy check, because accuracy was never the criterion doing the gatekeeping.

The US doesn't run this test at all

No CMS program, state licensing board, or federal agency currently requires a blinded accuracy test before an AI scribe vendor can sell to a US independent practice. HHS's Office for Civil Rights proposed an update to the HIPAA Security Rule in January 2025 that would have required practices to inventory the AI tools touching patient data and fold vendor risk into Business Associate Agreement reviews. That rule was originally targeted for a final version in May 2026. It has now slipped to a projected 2027 date, per the OMB's own regulatory tracking. The accuracy gap Ontario just documented in public has no US regulatory equivalent watching for it, formally or informally.

NHS England, by contrast, already requires more on paper: guidance issued in April 2025 directs NHS organizations to use only AI scribe systems that have been assessed for safety and compliance before deployment. Ontario's own auditors cited that NHS guidance directly, as an example of stronger oversight than the province's own process delivered. A US independent practice sits in the largest oversight gap of the three.

What Ontario's own findings say a practice should actually do

The most practical finding in the whole report has nothing to do with vendor scoring. Doctors using the approved AI scribes were never required to confirm, through a sign-off feature or any other mechanism, that they had reviewed the system-generated notes before those notes became part of the patient record. Guidelines existed. Attestation didn't.

That's the gap worth closing in your own practice regardless of which vendor you use or which country wrote the procurement rules. Four things follow directly from what this audit found:

Ask any AI scribe vendor for its own accuracy testing methodology and results, in writing, before you sign anything, not a marketing claim about hallucination rates. We've covered why vendor-reported hallucination rates and independent estimates rarely agree, and this is exactly why.

Build a mandatory human review step into your workflow for every AI-generated note, with an actual sign-off, not an assumption that "good enough most of the time" is good enough for a chart entry involving medication or mental health details.

Treat a vendor's compliance badge, HIPAA BAA, or "government-approved" claim as a floor, not a ceiling, on your own scrutiny. A signed BAA covers HIPAA. It doesn't cover accuracy, and as we've written before, it doesn't cover state consent law either.

If you're weighing a general-purpose AI tool against a purpose-built scribe, remember accuracy testing is uneven across both categories. Our existing comparison of general AI tools against dedicated scribes and our breakdown of bundled versus standalone AI scribe options for physio practices are both worth reading before you commit to either path.

The call

Twenty vendors went through formal government testing, and all twenty came out with a documented accuracy failure, because accuracy was worth 4% of the score and nothing enforced a minimum. If a government procurement process with real testing behind it produces that result, a sales demo with no testing behind it at all deserves more scrutiny, not less. Before you buy or renew an AI scribe contract, get the vendor's own accuracy data in writing, and put a real sign-off step in front of every note the AI generates. Ontario's auditors just showed what happens when nobody does either.

If you want an independent read on how your practice is actually verifying AI-generated clinical documentation today, versus how it should be, the AI Opportunity and Growth Assessment covers exactly this ground. Start with a free 20-minute discovery call.

A government test found errors in 100% of the AI scribes it reviewed. If you haven't asked your own vendor for its accuracy data yet, now is the time. Book a 20-minute call.

The Clinical AI Briefing

One practical AI insight for healthcare practices every week. No hype. Evidence and outcomes only.

Related: Dental AI scribes: the hallucination rate no vendor agrees on · Your AI scribe's BAA covers HIPAA. It doesn't cover this $5,000-a-patient lawsuit. · $20 vs $150: ChatGPT or an AI scribe