AI Dent
029 AI Radiograph Reading 1,244 words · 6 min

Auditing What Changed After You Switched On Radiograph AI

Every practice that adopts radiograph AI asks the same question before signing: does it find the caries? Almost nobody asks the one that matters a year later: what happened to our treatment thresholds once it started drawing boxes on the bitewings?

That is the argument of this post. An AI overlay does not just detect disease. It nudges clinicians, quietly and unevenly, towards treating more lesions earlier. If you do not measure that, you will not see it. A dental ai clinical audit, run before and after go-live, is the only way to find out whether your tool is improving care or simply raising your restoration rate.

Why thresholds drift

Tools such as Overjet, Pearl’s Second Opinion, Dentistry.AI, and Videa (now part of Pearl’s wider ecosystem in some markets) are trained to flag radiolucencies, including small enamel and outer-dentine lesions. Many of those lesions are, clinically, watch-and-review territory: a lesion confined to the outer half of dentine in a low-risk patient may well be managed with preventive care and a recall bitewing.

The software does not know your patient’s caries risk, diet, fluoride exposure or whether they turn up for recall. It marks the lesion. The clinician then sees a coloured outline on screen and has to actively decide not to act on it. Defensive practice does the rest. Nobody wants to explain to a complaints handler why a flagged lesion was left alone.

None of this is a fault in the product. It is a behavioural effect of putting a confident-looking prompt in front of a busy clinician. The published literature on computer-aided detection generally shows sensitivity going up, often with some loss of specificity. Higher sensitivity with lower specificity means more false positives, and in dentistry a false positive can end up as a filling.

The audit: what to compare

Keep it simple enough that a practice manager and one associate can run it in an afternoon per quarter. You need two periods of equal length.

  • Baseline: the 12 months before go-live (or the most recent 6 if your records are patchy).
  • Post-adoption: the same length of time, starting after a four-week bedding-in period.

Pull these measures from your practice management system (SOE, Dentally, Software of Excellence, R4 and Exact all support exports or reports for treatment codes):

  1. Restorations per 100 bitewing examinations. New direct restorations placed on posterior teeth, divided by the number of bitewing sets taken, times 100.
  2. Proportion of restorations on early lesions. Where your notes or charting record depth (E1/E2/D1 versus D2/D3), what share of restorations were on lesions limited to enamel or outer dentine?
  3. Re-treatment within 24 months. Replacement or repair of a restoration placed in the audit window, at the same surface, plus any extraction or endodontic treatment on a tooth restored in the window.
  4. Preventive actions logged. Fluoride varnish applications, high-fluoride toothpaste prescriptions and documented risk-based recall intervals per 100 patients.
  5. Clinician-level split. Everything above, per clinician, anonymised in the first pass.

A worked example

Take a fictional mixed NHS and private practice with four clinicians. Figures are illustrative, but the arithmetic is the kind you will actually do.

Measure12 months before12 months after
Bitewing sets taken1,8401,910
New posterior restorations612784
Restorations per 100 bitewing sets33.341.0
Restorations on enamel/outer-dentine lesions98 (16%)211 (27%)
Fluoride varnish applications540560

Restorations per 100 sets rose from 33.3 to 41.0, a jump of 7.7 points or about 23%. The number of bitewings barely moved, so this is not just more imaging. The share placed on early lesions rose from 16% to 27%. Varnish applications stayed flat, which suggests the extra disease was not being managed preventively.

Now the question the vendor cannot answer for you: was that extra 172 restorations caught disease, or premature intervention? The re-treatment figure helps. Suppose that, at 24 months, you can only follow the first-year cohort.

CohortRestorationsRe-treated at 24 monthsRate
Pre-adoption612559.0%
Post-adoption7849412.0%

A three-point rise in re-treatment is not proof of harm. But restorations placed on small lesions enter the repeat restorative cycle earlier, and a higher failure rate in a bigger cohort is exactly what an over-treatment drift looks like. If, instead, re-treatment had fallen while early detection rose, you could say with more confidence that the tool was catching lesions that would otherwise have progressed.

Reading the result honestly

A higher restoration rate is not automatically bad. Some practices were under-diagnosing. If your baseline was 18 restorations per 100 bitewings and your local peers sit near 30, an increase might be correction. So compare against something:

  • Your own pre-adoption figures (the main comparison).
  • The NHS BSA dental statistics for your contract area, which give a rough sense of the local picture.
  • Your clinicians’ individual baselines. One associate going from 28 to 52 per 100 sets while the others move by four points is a conversation, not a statistic.

Watch for confounders. A new associate, a change in patient mix after a contract variation, a shift in the proportion of private patients, or a hygienist-led recall drive can each move these numbers. Note them next to the table, in plain words.

What to do with the findings

If thresholds have drifted upwards, do not switch the tool off. Change how it is used.

  1. Set a written treatment threshold. For example: no operative intervention on lesions confined to enamel or the outer third of dentine in patients at low or moderate caries risk, unless there is cavitation on clinical examination. Align it with SDCEP’s Prevention and Management of Dental Caries in Children and the equivalent adult guidance.
  2. Record risk first. Make caries risk assessment a required field before a restorative plan can be saved.
  3. Review the sensitivity setting. Some platforms let you adjust which findings display. Ask your vendor in writing what the default operating point is and whether it can be changed.
  4. Re-run the audit at six months. Share clinician-level data openly, in a peer review meeting, not as a performance tool.

Log all of it. If the CQC asks how you assure yourself that new technology is not changing care in unintended ways, a two-page audit with dates, numbers and actions is a strong answer. A brochure is not.

Before you sign, not after

Add three lines to any contract discussion: that you will be given the tool’s sensitivity and specificity at the configured threshold on a population like yours, that you can export finding-level data (what was flagged, what was accepted or dismissed), and that you can change the display settings. Accept/dismiss data is gold. If clinicians accept 90% of flags on early lesions, you have learned something about threshold drift without waiting a year.

For a fuller look at how these tools are evaluated clinically, see our guide to AI radiograph reading.

Pull your baseline numbers this month. Even a rough count of restorations and bitewings from the last twelve months will give you something to compare against, and you can only get that baseline once: after go-live, it is gone.