Ask a dental principal what their AI scribe's accuracy rate is, and most will quote whatever number the vendor's sales page put in front of them. Ask a second principal using a different vendor, and the number can be ten times higher or lower, describing a completely different kind of error. Neither is lying. They're reading two different measurements off two different rulers, and nobody sold either of them the ruler.
That gap matters more in dentistry than in most specialties. A hallucinated line in a periodontal chart isn't just an embarrassing typo, it's a clinical record that can misstate pocket depths, mobility scores or restorative findings a patient's future care depends on, and a record an insurer or the GDC can later ask you to defend.
What an AI scribe is actually doing, step by step
Strip away the marketing and an ambient AI scribe is a three-stage pipeline. First, automatic speech recognition (ASR) turns the audio of the appointment into a time-stamped transcript. Second, a natural language processing layer picks out which parts of that transcript are clinically relevant and who said what. Third, a large language model restructures that material into a note, typically formatted as a SOAP note or, for dentistry, mapped against a chart template and CDT procedure codes (Suki, "What Is Ambient Clinical Intelligence? ACI Guide for 2026").
The first stage has genuinely improved. Suki's 2026 guide puts state-of-the-art clinical ASR at well under 10% word error rate on clean audio, down from roughly 50% a decade ago, a vendor-reported figure but consistent with the wider direction of published speech recognition research. That improvement is real. It's also not the part causing the current safety debate.
So what for you: transcription accuracy was the old problem and it's mostly solved. The current risk sits one stage further down the pipeline, in what the model does with an accurate transcript once it starts writing.
Where hallucination actually happens
A hallucination is the model confidently writing something into the note that the encounter didn't contain: a finding that wasn't examined, a symptom the patient never mentioned, in the worst documented cases, an entire section of an exam that never took place. A 2025 comment piece in npj Digital Medicine, from researchers at Columbia University's Data Science Institute, names hallucination as the single most important clinical safety consideration in ambient scribing, and reports that some systems have produced full physical-examination narratives for examinations that were never performed (Zhang et al., "Beyond human ears: navigating the uncharted risks of AI scribes in clinical practice," npj Digital Medicine, published online 24 September 2025).
That's a peer-reviewed, non-vendor source, and it's the reason "how accurate is the transcript" is the wrong question to ask a vendor. The transcript can be near-perfect and the generated note can still contain fabricated clinical content, because the third stage of the pipeline is generating language, not just repeating what it heard.
So what for you: when you evaluate a scribe, ask specifically about hallucination handling, not word error rate. A vendor who only talks about transcription accuracy is answering a question from a decade ago.
The number nobody agrees on
This is where the two figures in the FAQ above come from, and why they don't contradict each other despite looking like they should. Suki's 2026 guide, a vendor source and worth reading with that in mind, reports leading systems at a 1-3% hallucination rate on clinical content specifically, meaning roughly one to three fabricated clinical facts per hundred notes. A separate composite drawn from published evaluations and clinical-notes compliance guides, not independently verified and flagged as such in our own research file, puts the share of notes containing at least one hallucination of any kind, including smaller omissions and misattributions, at one in five to one in three.
Read literally, a practice could see both figures describing the exact same software. One counts severity-weighted fabricated facts as a percentage of total content. The other counts any note with at least one flaw, of any size, as a percentage of notes. A note can pass the first test and fail the second.
So what for you: the question to put to any vendor isn't "what's your hallucination rate", it's "what exactly does that number count, per note or per fact, and against what independent evaluation, not just your own testing." Most sales conversations never get asked that, and most vendors won't offer the distinction unprompted.
Why dentistry carries a sharper version of this risk
Dentists spend an estimated 10 hours a week on clinical documentation, charting tooth-specific findings, periodontal measurements, treatment narratives and CDT coding, according to vendor-adjacent industry estimates from AI charting providers (DeepCura; Twofold, 2026 dental AI guides), a figure worth treating as directional rather than precise given its source. That workload is exactly why voice-driven charting tools like Denti.AI's and DentScribe's AI voice perio charting are attractive: speaking findings aloud during an exam is faster than typing them between patients.
It's also why an error matters more here than in a general medical note. A hallucinated pocket depth or a misattributed mobility score doesn't just read badly, it becomes the baseline the next hygiene visit and the next periodontal assessment gets compared against. A CDT code entered incorrectly by an AI charting assistant is a billing and claims problem as well as a clinical one. The stakes of an unreviewed hallucination compound every time the record is reused, which in dentistry is constantly.
So what for you: treat periodontal charting and CDT coding as the two highest-scrutiny outputs from any AI scribe or charting tool, not general narrative notes, because errors there propagate into future clinical decisions and into claims.
What GDC and BDA already expect you to do about it
Neither the General Dental Council nor the British Dental Association has published a hallucination-specific checking protocol. What already exists is broader, but it still lands squarely on this problem. The GDC's Standard 3.1 requires valid, informed consent before treatment, a standard our earlier analysis of GDC versus BDA AI guidance covers in full, and the BDA's May 2026 advice states plainly that professional accountability for any clinical decision stays with the dentist, comparing AI to loupes: a tool that enhances what you do, not one that does it for you (British Dental Association, "Should dental practices use AI?", 19 May 2026).
Put together, that means every AI-generated note needs a clinician's review before it becomes part of the permanent record, the same discipline you'd apply to a note drafted by a new associate or a locum. CQC folds AI-specific governance checks into its well-led framework for practices registered in England, which our earlier piece on CQC's 2026 well-led rewrite sets out, and an unreviewed hallucination sitting in a patient record is exactly the kind of gap that framework is designed to catch.
So what for you: "the AI wrote it" is never a defensible answer to why a chart entry is wrong. Build a review step into the workflow, not as a formality, but as the actual safeguard against the failure mode this article describes.
What this means for your practice
Three actions, in order of priority.
First: before signing with any AI scribe or charting vendor, ask them directly what their quoted accuracy or hallucination figure counts, per-note or per-fact, and whether it comes from independent evaluation or their own internal testing. Get the answer in writing, not just in a sales call.
Second: build a mandatory review step for every AI-generated chart entry before it's finalised, with periodontal measurements and CDT codes checked first, given how far those errors propagate through future visits and claims.
Third: don't let a headline accuracy percentage substitute for your own short pilot. Run any new scribe or charting tool against a handful of real appointments and compare its output line by line against what actually happened before rolling it out practice-wide. If you want an independent, vendor-neutral read on which tools are worth that pilot for your specific practice, our AI Opportunity & Growth Assessment™ scores exactly this, and you can book a 20-minute call to talk through where your practice sits before committing to anything.
The single most important thing to take from this
Speech recognition stopped being the hard problem years ago. What a large language model chooses to write once it has an accurate transcript is the part still being figured out, by academics, by vendors, and increasingly by regulators watching NHS-scale rollouts of the same underlying technology. Practices that ask vendors to define their own accuracy numbers, rather than accepting a percentage at face value, are the ones who'll catch a hallucinated finding before it becomes next year's baseline instead of after.
The Clinical AI Briefing
One practical AI insight for healthcare practices every week. No hype. Evidence and outcomes only.