When the AI and the Dentist Disagree: Building a House Rule
Two associates. Same practice, same Pearl Second Opinion licence, same Tuesday afternoon. Both open a set of four bitewings. Both see the software throw a proximal caries box on the distal of an upper right six. One of them charts it as caries into dentine, books a 3 UDA Band 2 and tells the patient the AI picked something up that was hard to see. The other decides it is cervical burnout dressed up as a lesion, dismisses the overlay, writes nothing about it in the notes, and moves on.
Neither of them is wrong, exactly. That is the problem.
The inconsistency you cannot see from the front desk
Ad hoc resolution feels harmless because every individual decision is defensible. A clinician looked at a radiograph, applied judgement, made a call. That is the job. But scale it: a three surgery practice seeing 1,400 courses of treatment a quarter, four clinicians including a Thursday locum, and a diagnostic aid that flags something on roughly one in three bitewing sets. You now have four private conventions operating under one practice name, and nothing in the record that distinguishes “the AI was wrong” from “the clinician never looked at the AI”.
Pull the notes and you will find it. Here is the pattern from a practice that ran this audit properly: 200 consecutive bitewing sets across four clinicians over eight weeks, counting how often an AI-flagged proximal lesion that the clinician did not restore was mentioned anywhere in the written record.
| Clinician | Sets reviewed | AI proximal flags | Flags addressed in notes | Rate |
|---|---|---|---|---|
| Principal | 62 | 71 | 58 | 82% |
| Associate A | 55 | 64 | 9 | 14% |
| Associate B | 48 | 59 | 41 | 69% |
| Locum (rotating) | 35 | 43 | 2 | 5% |
Same software. Same patients, near enough. A 77 point spread in documentation behaviour, and it is invisible until someone counts.
That spread does three things to you. It makes your charting data useless for anything downstream, so you cannot compare restoration rates between associates, cannot audit your own diagnostic thresholds, cannot answer a CQC inspector who asks how you use the tool. It creates a complaints exposure: when a patient who saw Associate A in March returns in November and Associate B restores a lesion the software flagged both times, the record shows a lesion that appeared from nowhere. And it quietly hands clinical policy to whoever configured the software, because in the absence of a practice rule, the vendor’s default sensitivity setting becomes your treatment threshold.
What “disagreement” actually means
Before you can write a rule you have to stop treating this as one phenomenon. Every AI second opinion dental radiograph workflow produces at least four distinct kinds of conflict, and they need different handling.
Depth disagreement. You and the software both see a lesion, you grade it differently. The AI calls D1, you call E2. This is the common case and by far the most consequential, because the enamel/dentine line is where restoration decisions live.
Existence disagreement. The software flags a surface you read as sound, or you see something it missed. Overlapping contacts, restoration margins, radiolucent bases under old amalgams and cervical burnout on premolars generate most of the false positives you will meet.
Measurement disagreement. Overjet will return a CEJ to bone crest distance in millimetres. If it reports 3.4 mm distal of an LR6 and your periodontal charting says 4 mm pocket with no recession, the numbers are not in conflict, but the story you tell the patient might be.
Scope disagreement. The software reports a finding outside what you were looking at. Diagnocat on a CBCT volume taken for implant planning will report on every tooth in the field. You now own that finding whether you wanted it or not, and IR(ME)R 2017 requires a recorded clinical evaluation of the exposure regardless.
Lump these together and your policy will be vague. Separate them and it writes itself.
The house rule
The rule has two halves: a fixed set of dispositions, and a short vocabulary of reasons. Nothing else. Resist the urge to build a decision tree.
Every AI finding gets exactly one of four dispositions, recorded with a code:
| Code | Disposition | What it means | Note requirement |
|---|---|---|---|
| A | Accept | Clinician agrees with the finding and its grade | Chart as normal, no extra note |
| O | Override | Clinician rejects the finding | Reason code mandatory |
| M | Monitor | Finding accepted, intervention not indicated | Review interval stated |
| E | Escalate | Unresolved after review | Named second opinion or further imaging |
Override reasons are a closed list of six. A closed list is the entire point: free text produces prose that nobody can audit.
- O1 Image artefact (burnout, cone cut, processing)
- O2 Overlapping contact, surface not assessable
- O3 Existing restoration or base mimicking lesion
- O4 Clinical examination contradicts (visual, tactile, transillumination)
- O5 Prior radiograph shows no progression
- O6 Anatomical variation
Now the part that stops most of the inconsistency in one line. Set a house threshold that removes clinician discretion from the enamel range: no AI finding graded at enamel depth, at any confidence level, triggers an operative note. It gets disposition M with a stated interval, and prevention. This is consistent with SDCEP and with NICE recall guidance, it is what most of your clinicians already believe, and writing it down means Associate A and Associate B produce the same chart on the same tooth.
Worked example: the UR6 distal
Back to Tuesday. Under the house rule, both associates now do this:
14/10/2026 BW x4, FGDP quality diagnostically acceptable.
AI (Pearl Second Opinion v3): proximal radiolucency UR6 D,
reported dentine, outer third.
Clinician read: E2, confined to enamel. Marginal ridge intact,
no cavitation on direct vision, ICDAS 0.
Disposition: O (override) + M (monitor)
Reason: O4 clinical exam contradicts + O5 stable vs BW 09/03/2025
Plan: 5000ppm F toothpaste, diet advice re. squash, BW 12/12 months.
Discussed AI finding with pt, explained why not restoring today.
Fifty seconds of typing. What it buys you: the next clinician sees the lesion was seen, graded, and deliberately not restored. The patient was told. If it progresses, the progression is documented against a baseline. If a complaint lands in 2029, the note answers the question before it is asked.
Compare that to the version where the flag was silently dismissed. There is no record that the surface was ever assessed, and the defence becomes an argument about what a clinician probably did three years ago.
Worked example: the locum who has never used it
Your Thursday locum arrives, logs into Dentally, opens Romexis, and the AI overlay appears on a radiograph. They have used VideaHealth at one practice and nothing at another. Without a written rule they will do what they did last week somewhere else.
Give them one side of A4 in the induction pack. Not training on the software, which the vendor handles. A statement of what this practice does with the output:
AI RADIOGRAPH FINDINGS: PRACTICE POLICY
Version 2.1 | Approved: 01/06/2026 | Review: 01/06/2027
Applies to: all clinicians incl. locums, therapists, hygienists
1. The AI is a decision support tool (UKCA marked, Class IIa).
The clinical decision and the IR(ME)R clinical evaluation
remain with the treating clinician. Always.
2. Every AI finding on every radiograph receives a disposition:
A / O / M / E. No finding is left unaddressed.
3. Overrides require a reason code from the list of six.
4. Enamel-depth findings are never restored on radiograph alone.
Disposition M, interval stated, prevention recorded.
5. Escalation (E) goes to the on-site principal same day, or
to a second set of images at the next appropriate visit.
Never to the patient as "the computer thinks".
6. Discrepancies are reviewed monthly. Sample of 20.
Six clauses. A locum reads it in ninety seconds and charts like everyone else.
Running it, and the numbers that force your hand
The monthly review is where a policy stops being a document and becomes a practice norm. Twenty cases, ten minutes at the staff meeting, pulled at random from SOE Exact or whatever you run. You are looking for two things: overrides with no reason code, and reason codes that cluster oddly. If one clinician’s overrides are 80% O1 image artefact, either your X-ray technique needs attention or that clinician is using O1 as a shrug.
Specificity is the number that should shape your threshold, not sensitivity. Vendor validation work tends to lead with sensitivity against expert consensus, which is the flattering figure. Think about what happens downstream instead. A four bitewing set exposes somewhere around 32 to 38 assessable proximal surfaces. At 95% specificity on genuinely sound surfaces, you get roughly 1.7 false flags per patient. Drop to 90% and it is 3.5. A practice taking 1,600 bitewing sets a year is therefore fielding somewhere between 2,700 and 5,600 false positives annually, and every one of them is a moment where an associate either writes something down or does not.
Set against that, the per case cost of the rule is about a minute. If you are still deciding which system to buy, or whether the detection thresholds are adjustable at practice level, the AI radiograph reading guide covers what to ask vendors before you sign.
The objection worth answering
Someone will say this is defensive charting, that it adds clicks, that clinical judgement should not need a form. Fair enough as a sentiment. But the alternative on offer is not unfettered judgement, it is four different unfettered judgements producing four different records of the same mouth, and a practice that cannot describe its own diagnostic standard when asked.
A house rule does not tell a clinician what to think about a radiolucency. It tells them what to write down once they have thought about it. Those are very different things, and only one of them survives contact with an associate who leaves in eighteen months and takes their conventions with them.
Print the six clauses. Put a version number and a review date on it. Get every clinician who reads a radiograph in your building to sign the bottom, including the ones who only come on Thursdays.