AI Dent
005 Clinical Notes and Charting 1,818 words · 8 min

What an AI Dental Scribe Actually Captures During an Exam

Ambient scribes are being sold into UK practices on a single promise: you talk, it writes, you sign. That promise is roughly 70% true, and the missing 30% is the part with your GDC number attached to it.

So rather than argue about it in the abstract, here is a real check-up, start to finish, with the raw capture and the generated note side by side. The transcript below is a composite drawn from routine Band 1 exams, lightly anonymised, but nothing has been cleaned up to make the AI look better or worse than it is. If you are evaluating ai dental note taking for an NHS or mixed list, this is the level of detail you need before you sign anything.

The eleven minutes

Patient is 54, six-monthly recall, complaining of cold sensitivity lower left. Nurse present. Microphone is a phone on the bracket table, which is how most pilots actually run.

[00:00] D:  Morning Janet, come through. How've you been?
[00:05] P:  Not bad. That tooth at the back on the left's been funny with cold.
[00:11] D:  The one we looked at in March?
[00:13] P:  I think so. Mainly water. It goes off after a few seconds.
[00:19] D:  Any pain at night? Keeping you awake?
[00:21] P:  No, nothing like that.
[00:24] D:  Good. Lie back for me. [chair motor] Sharon, light please.
[00:31] D:  [suction] Upper right, seven's got the MOD amalgam, margins look
            alright. Six, occlusal composite, sound. Five, four, three sound.
[00:52] D:  Upper left, six is the crown, PFM, margin's fine. Seven... seven's
            got a bit of wear on the palatal.
[01:14] D:  Lower left now. Six, distal amalgam, and there's... hmm. There's
            some shadowing under that distal margin. Could be shine-through,
            could be nothing. [suction, 4 sec]
[01:33] D:  Is it this one, Janet? [tap] That one?
[01:36] P:  That's the one. That's cracked, I'm sure it's cracked.
[01:40] D:  Let's not jump ahead. Cold test, Sharon.
[01:58] D:  Normal response, settles. No lingering.
[02:20] D:  BPE. Two, two, two. Three, two, two.
[02:41] D:  Let's take a left bitewing. Nothing since 2024, is that right?
[02:46] N:  Last bitewings were the 14th of March 2024.
[09:10] D:  Right. There is something distal on the six but it's not through
            to dentine that I can see. I'd rather monitor it and get you back
            in three months than start drilling today.
[09:30] P:  So you're not filling it?
[09:32] D:  Not today. Fluoride varnish, and I want you using a 5000 paste at
            night. Sharon can sort the prescription.
[10:40] D:  Six months for the check, three months for that tooth.

What the scribe got right

I ran this through a general-purpose ambient scribe of the kind now common in UK primary care (Heidi and Tandem are the two most practices seem to be trialling, with Tortus appearing in NHS trust settings), plus Bola AI, which is built for dental and does voice-driven perio and hard tissue charting rather than narrative notes.

The generic scribes were genuinely good at four things:

The history. Cold sensitivity, lower left, water, non-lingering, no nocturnal pain. All four attributes captured, correctly attributed to the patient, correctly placed under a history heading. That is a proper SOCRATES-shaped symptom history assembled from a conversation that never once used the word “history”.

The advice and the plan. Fluoride varnish, 5000 ppm paste at night, three-month review of the LL6, six-month recall. Nothing dropped. This is the part that indemnity organisations write case reports about, and it is the part hand-typed notes lose at 5:40pm on a Friday.

The negotiation. “So you’re not filling it?” followed by “Not today” was rendered as a line about the monitoring decision being discussed with the patient. That is a consent breadcrumb you would probably not have typed yourself.

Speed. Note drafted in about 20 seconds after the recording stopped. Budget 45 to 90 seconds of review per patient and you are still well ahead of the two to four minutes most of us spend typing a Band 1.

What it invented

Now the other column. Every item here appeared in at least one of the generated notes.

Tooth numbering drifted. The dentist said “seven’s got the MOD amalgam” in the upper right quadrant. One scribe wrote UR7 MOD amalgam. Another wrote 17 MOD amalgam, silently converting to FDI. A third wrote UR7 MO amalgam, dropping a surface. Audio of “MOD” through a mask, with suction running, is genuinely marginal: a high-speed handpiece puts 70 to 95 dB at the operator’s ear, and surface letters are short, unstressed and acoustically similar. The scribe does not know it got it wrong, so it does not flag it.

The BPE scrambled. Spoken as “two, two, two, three, two, two”. One output rendered the sextants as 2 2 2 / 2 2 3. The lone 3 moved from lower right to lower left. Six digits, one transposition, and now your periodontal diagnosis and your next appointment length are both built on a wrong sextant.

A crack appeared in the clinical findings. The patient said “that’s cracked, I’m sure it’s cracked”. The dentist said “let’s not jump ahead”. One note contained the line Suspected cracked cusp LL6. Nobody clinically suspected a cracked cusp. The model saw a dental word near a symptom and promoted it.

Findings that were never examined got documented as normal. Two of the three drafts included some version of Soft tissues NAD. Oral cancer screening performed, no abnormality detected. Listen to the transcript again. There is no soft tissue exam in it. The model produced that line because in the enormous corpus of dental notes it has seen, that sentence follows this kind of paragraph. This is the single most dangerous behaviour in ai dental note taking, because the fabrication is always a normal finding, always plausible, and reads exactly like something you would write.

The hedge got upgraded. “Could be shine-through, could be nothing” became Early distal caries LL6 in one draft. A monitoring decision was quietly rewritten as a diagnosis you did not make and would not defend.

None of this is a bug that gets patched next quarter. The underlying speech models are the same family as OpenAI’s Whisper, and a 2024 FAccT paper (“Careless Whisper”, Koenecke et al.) found hallucinated content in roughly 1% of transcribed segments, with around 38% of those hallucinations containing material that could cause harm if believed. Silence and noise are the triggers. A dental surgery supplies both, constantly.

The diagnosis line is not a transcription problem

Here is the structural reason the last line stays yours, and it has nothing to do with how good the model gets.

Under IR(ME)R 2017, the bitewing in that appointment had to be justified by a practitioner before exposure and reported afterwards. Justification is a judgement about this patient, this symptom, this prior exposure date. A scribe can transcribe “let’s take a left bitewing”. It cannot justify it, and if it writes a justification sentence, that sentence is fabricated evidence in a regulated process. Same for the radiographic report: a grade 1 with a written finding is an act of clinical reading, not an act of listening.

GDC Standard 4.1.1 requires complete and accurate records, and the CGDent/FGDP record-keeping guidance expects the diagnosis and the reasoning behind the treatment decision to be recorded. When the DDU or Dental Protection open a file three years from now, the question will be whether your note shows that you thought about it. A sentence generated from a genre template does not show that.

Then there is the money. A Band 1 course of treatment claims one UDA, worth somewhere between about £25 and £35 depending on your contract value. The FP17 is a claim made by you, on your performer number, about clinical care you provided. If the note underneath it contains an oral cancer screening that never happened, that is not a documentation slip.

Practices running this well treat the scribe as a stenographer with no clinical authority. It drafts history, exam narrative, advice, plan. The clinician writes or confirms three things by hand: the tooth numbers with surfaces, the BPE, and the diagnosis and reasoning line. Our full breakdown of how that split works in a mixed NHS list sits in Clinical Notes and Charting.

Running a pilot that actually tells you something

Pick 20 consecutive check-ups. Record and keep your own hand-typed note alongside the generated one. Then score every draft against five columns:

CheckWhat you are countingPass mark
Tooth notationWrong quadrant, wrong number, wrong or dropped surface0 errors in 20
PerioBPE digits correct and in the right sextant order0 errors in 20
Fabricated normalsFindings documented that were never examined0 in 20
Confidence driftHedge rendered as diagnosis0 in 20
AttributionPatient’s words stated as clinician’s findings0 in 20

Twenty is not a huge sample, and that is the point: if you find two tooth-numbering errors in twenty routine exams, you have your answer without needing statistics. One error per ten patients is 400-odd wrong notes a year on a single full list.

Three configuration changes make the difference between a tool you trust and one you unpick:

Turn off template completion. Every vendor has a setting, sometimes called “comprehensive note” or “auto-complete sections”, that fills unstated sections with defaults. Off. You want visible gaps, because a gap prompts you and a plausible sentence does not.

Do your charting separately. Voice-driven charting into Dentally, SOE Exact, Carestream R4 or Systems for Dentists is a different task with a constrained vocabulary and a confirmation loop, and the dedicated tools (Bola AI being the obvious one) are better at it than a narrative scribe will ever be. Keep radiograph AI separate too: Pearl’s Second Opinion, Overjet and VideaHealth read images, they do not listen, and their outputs still need your report.

Settle the data question before the clinical one. Ambient capture of a consultation is special category health data under UK GDPR Article 9. You need a DPIA, a signed data processing agreement, clarity on whether audio leaves the UK and how long it is retained, and, if you hold an NHS contract, a DSPT submission that reflects what you are actually doing. Several practices have run three-month pilots and then discovered the retention terms were unacceptable.

The honest position after eleven minutes of audio is that the machine heard the appointment well and understood none of it. It got Janet’s symptom history better than I would have typed it, and it put a crack in her tooth that only she believed in.