How Accurate Is AI Caries Detection on Bitewings?
Ask a rep how accurate their caries AI is and you will usually get one number, somewhere between 90% and 98%, delivered with confidence. That number is almost always useless to you. It might be an area under the ROC curve from a research dataset, or a raw agreement rate on a set of images where most surfaces were obviously sound, or a figure calculated only on frank dentine lesions where any competent dentist would have got it right anyway.
This page takes ai caries detection accuracy apart properly: what the published figures actually say, why sensitivity and specificity pull in opposite directions, what the numbers translate to across a normal day of bitewings in an NHS or mixed practice, and how to bench-test a system on your own images before you sign anything. If you want the wider picture first, including how these tools handle bone levels, periapical pathology and workflow integration, the AI radiograph reading overview covers the category as a whole.
The single accuracy number is the wrong number
Caries detection on a bitewing is a binary call repeated across every proximal surface in the image. Four outcomes are possible: the software flags a lesion that is really there (true positive), flags one that is not (false positive), stays quiet on a real lesion (false negative), or correctly stays quiet on sound enamel (true negative).
“Accuracy” lumps all four together, which flatters any model working on a population where disease is rare. Imagine a set of bitewings with 32 scoreable proximal surfaces and two real lesions. A model that flagged nothing at all would score 30/32, or 94% accuracy, while missing 100% of the disease. That is the arithmetic behind a lot of impressive-sounding marketing.
The two figures you need are sensitivity (of the lesions that are there, what proportion does it flag?) and specificity (of the sound surfaces, what proportion does it leave alone?). They trade against each other. Every one of these products has an internal confidence threshold, and most let you move it. Push the threshold down to catch more early lesions and you will catch more noise with them.
What the published UK evidence actually shows
The most directly relevant British study is the ADEPT trial, published in the British Dental Journal in 2021, which tested Manchester Imaging’s AssistDent on enamel-only proximal caries. Unaided, the dentists in that study detected around 24% of enamel-only lesions at a specificity of roughly 95%. With the AI overlay, sensitivity roughly doubled to about 46%, and specificity dropped to around 87%.
Read those four numbers together rather than picking the flattering pair. The software found nearly twice as many early lesions. It also roughly tripled the false positive rate, from about 1 in 20 sound surfaces to about 1 in 8. Whether that is a good trade depends entirely on what you do next with a flagged surface, and we will come back to that.
Broadly similar patterns appear in the wider literature. Meta-analyses of unaided bitewing reading have put pooled sensitivity for proximal lesions across all depths in the region of 0.24 to 0.45, rising substantially once a lesion is clearly into dentine, with specificity typically above 0.90. Randomised and reader-study work on AI assistance from groups in Berlin and elsewhere tends to land in the same place: a real gain in sensitivity, concentrated in shallow lesions, bought with some loss of specificity.
Vendor-reported figures sit higher. Pearl’s Second Opinion, Overjet and VideaHealth have all published or submitted reader studies supporting their FDA clearances, and the headline sensitivities quoted in those materials commonly sit in the 0.80 to 0.90 range. Those studies are not fraudulent, but they are run on curated image sets, with a defined reference standard, often on sensor hardware the model was trained on. Treat them as an upper bound, not an expectation.
The prevalence trap: what a flag is worth in your practice
Sensitivity and specificity are properties of the model. Positive predictive value, which is the thing you actually experience, depends on how much caries your patient list has.
Take a stable, largely preventive adult list where about 5% of scoreable proximal surfaces carry a radiographic lesion. Run a model at sensitivity 0.85 and specificity 0.90 across 1,000 surfaces:
| Lesion present | Lesion absent | Total | |
|---|---|---|---|
| AI flags | 42.5 | 95 | 137.5 |
| AI silent | 7.5 | 855 | 862.5 |
| Total | 50 | 950 | 1,000 |
Positive predictive value comes out at 42.5 / 137.5, which is 31%. Roughly two out of every three boxes on the screen are on sound enamel. At around 32 scoreable surfaces per adult bitewing set, that is about three spurious flags per patient, every patient, all day.
Now hold sensitivity at 0.85 and lift specificity to 0.97. False positives fall from 95 to 28.5, and PPV rises to about 60%. Specificity is doing far more work here than sensitivity, and it is the number vendors are least keen to lead with.
Flip the population. On a high-needs NHS list where 20% of proximal surfaces carry a lesion, the same model at 0.85/0.90 gives 170 true positives against 80 false positives, a PPV of 68%. Identical software, wildly different lived experience. This is why a colleague’s enthusiastic recommendation transfers badly between practices.
What the output actually looks like
A typical per-surface output, before the interface prettifies it, reads something like this:
Patient 41822 Right posterior bitewing 2026-09-14
Finding Site Confidence
Proximal radiolucency, dentine UR6 mesial 0.88
Proximal radiolucency, enamel UR5 distal 0.61
Restoration margin discrepancy LR7 mesial 0.72
Proximal radiolucency, enamel LR6 distal 0.44 (below display threshold 0.50)
That 0.44 finding is the interesting one. It exists, the model saw something, and a threshold decision hid it from you. Ask any vendor whether you can see sub-threshold findings on demand, and whether the threshold is set per practice, per clinician or centrally by them. Some systems ship with a single fixed threshold tuned for the US insurance market, where a flagged lesion supports a claim, and that tuning is not neutral for a UK preventive workflow.
Where these systems are genuinely strong, and where they fall over
Frank dentine lesions are detected reliably, in the high 80s or better. That is also where unaided human sensitivity is already decent, so the marginal gain is modest.
Early proximal enamel lesions are where the real uplift sits, and also where the false positives cluster. Cervical burnout at the CEJ produces a radiolucency that looks convincingly carious to a model trained on annotated images, because it looks convincingly carious to dentists too. Overlapping contacts create apparent radiolucency where two enamel surfaces superimpose. Mach band effect at a restoration margin does the same.
Secondary caries around existing restorations is the weakest area across every product I have seen tested. Radiolucent cements and liners, composite with low radiopacity, and genuine margin gaps all present similarly. Expect materially worse specificity here and check whether the vendor reports it separately. Many report only “caries” as one pooled category, which hides the problem.
Occlusal caries barely registers. Bitewings are a poor tool for occlusal detection in the first place, and no amount of model architecture fixes an imaging limitation.
The ground truth problem nobody raises in the demo
Almost every published accuracy figure for dental caries AI uses expert consensus on the same radiographs as its reference standard. Three experienced dentists look at the image, agree what is there, and that becomes truth.
This is circular. A model trained and evaluated that way can only ever learn to agree with dentists reading radiographs, including their systematic errors. The handful of studies using histological validation on extracted teeth, or micro-CT, consistently report lower numbers than the consensus-based ones, because the harder reference standard exposes lesions nobody could see on film.
When a vendor quotes you a sensitivity, ask one question: against what reference standard? If the answer is expert consensus, the figure tells you how human-like the model is, not how correct it is.
Image conditions change the answer
These models are sensitive to the hardware that produced the image. A network trained largely on solid-state sensor images from one manufacturer will typically lose a few points of accuracy on phosphor plate images with different noise characteristics, and more again on scanned film. Exposure factors, plate wear, processing artefacts and horizontal angulation all shift performance.
Ask which sensors and which capture software appear in the training data. If you run Dürr PSP plates and the vendor trained mostly on Planmeca solid-state, that is not disqualifying, but it is a reason to weight your own bench test heavily over their published figures.
A bench test you can run in an afternoon
Pull 50 archived bitewing sets where you know what happened next, ideally cases where a surface was subsequently opened, restored or monitored to a documented endpoint. Run them through the trial version.
Record, per surface: did the AI flag it, did the original reporting clinician flag it, and what was actually found clinically. You will end up with four buckets. The ones that matter are the surfaces where AI and clinician disagreed. Sort them and read them back as a group.
Count the false positives per patient, not as a percentage. Three per set is tolerable. Eight is a system nobody will still be using in six weeks, because clinicians start clicking through the overlay without reading it, which leaves you worse off than having no AI at all.
Then repeat one thing: take ten patients who have had two bitewing sets at different visits with no intervening change, and check whether the AI flags the same surfaces both times. Repeatability is rarely reported and matters enormously if you intend to use these outputs to track lesion progression.
Regulatory and clinical governance points specific to the UK
Caries detection software is a medical device. In Great Britain it needs UKCA marking or an accepted CE mark under the current transitional arrangements, and it will normally be Class IIa. Ask for the certificate and the intended use statement, and read the intended use carefully: several products are cleared as an aid to detection, explicitly not as a diagnostic, and that wording defines where your liability sits.
IR(ME)R 2017 is unchanged by any of this. The practitioner justifies the exposure, and an AI finding on an existing radiograph can never justify taking another one. Selection criteria still govern radiograph frequency.
Most importantly for an NHS contract: a correctly identified enamel lesion is not an indication to restore. Delivering Better Oral Health is clear that early proximal lesions are managed preventively. If your AI doubles enamel lesion detection and your restorative rate follows it upward, you have not improved diagnosis, you have automated overtreatment, and that is a conversation you will eventually have with someone holding your UDA data.
Set the threshold, agree in a practice meeting what happens to a flagged E1 or E2 surface, write it down, and re-audit your restorative rate against the baseline three months after go-live.