How to Read a Dental AI Vendor’s Clinical Evidence Pack
The rep leaves a 24-page PDF. There is an AUC of 0.97 on page 6, a CE mark on page 11, a professor’s endorsement on page 19, and a slide showing dentists finding 30% more caries. None of it tells you the one thing you actually need to know: whether this software will work on your bitewings, captured on your phosphor plates, read by your associate at 4:40pm on a wet Tuesday in a practice where two thirds of the adult list is Band 1 and hasn’t been seen for three years.
That gap is not an accident, and it is not usually fraud. It is what happens when a product is validated for a regulator and then sold to a market the validation never touched. Dental AI clinical evidence packs are built to satisfy the US FDA’s 510(k) substantial-equivalence route or an EU MDR conformity assessment. Neither of those asks “does it help a UK GDP?” So you have to ask, and there are three places to look where the answer is almost always thinnest.
Weak spot one: where the images came from
Start at the dataset, not the accuracy figure. A deep learning model for radiograph reading is a compressed summary of the images it was trained on, and dental radiographs vary enormously by the kit that produced them.
Here is the concrete UK problem. A large share of general practice in this country still images with phosphor storage plates: VistaScan, Digora, CRx. Plenty of practices run older direct sensors. US dental support organisations, which is where most of the big training corpora were assembled, are heavily standardised on current-generation direct digital sensors with their own proprietary processing. Grain, contrast curve, sharpening, bit depth, artefact pattern: all different. A model that has never seen a PSP plate scan with a light crease in it has no idea what that crease is.
So the question is not “how many images was it trained on?” Every vendor will happily say millions. The questions that produce useful answers are narrower:
- How many images in the training set were phosphor plate, and from which scanners?
- How many came from UK or Irish practices, and how many from NHS general practice specifically as opposed to private or specialist settings?
- What was excluded? Edentulous arches, implants, heavy amalgam, paediatric mixed dentition, orthodontic appliances, and anything the annotators found ambiguous are the usual quiet exclusions.
- Was the test set collected at different sites from the training set, or split randomly from the same pool?
That last one matters more than any other line in the pack. A random split of one big pool tells you the model learned that pool. Krois and colleagues at Charité in Berlin published exactly this in Scientific Reports in 2021: dental image models tested on data from a centre they had not trained on lost meaningful accuracy, in the region of ten to fifteen percentage points depending on the task. If a pack reports one headline figure and no external validation, treat the figure as an internal engineering benchmark. It is a useful sign the code works. It is not evidence about your practice.
Pearl’s Second Opinion has been through FDA clearance for detection on adult periapicals and bitewings. Overjet holds multiple clearances covering caries and bone level measurement. VideaHealth has clearance for its caries detection. All three are real regulatory achievements and all three are compatible with a pack that never shows you a single UK image. Ask for the site list. If the answer is “we can’t share that for commercial reasons,” you have learned something.
Weak spot two: how the reader study was built
The reader study is the part of the pack that looks most like clinical evidence and is most often designed to flatter. What you want is a multi-reader multi-case study where each reader sees each case twice, once aided and once unaided, with a washout period between, and with the order counterbalanced. What you frequently get is a single-arm study where the AI is compared against a reference standard and the humans appear only as the people who built that standard.
Four design details change the result more than the algorithm does.
Disease prevalence in the test set. Research datasets are usually enriched, sometimes to 40 or 50% positives, because that produces tight confidence intervals with fewer cases. Your bitewings are not 50% positive for enamel-only proximal lesions. Enrichment leaves sensitivity and specificity intact but wrecks positive predictive value, which is the number your associate actually experiences as trust or irritation.
The reference standard. Histology is the honest answer and almost nobody has it, because you cannot section a tooth you are trying to keep. Most packs use consensus of two or three expert readers, sometimes with the AI’s output visible to them. If the ground truth is expert opinion, the ceiling of the study is expert opinion, and the model can be “better than dentists” while being worse than the truth.
Reading conditions. Calibrated diagnostic monitor in a darkened room, unlimited time, no patient in the chair, no history, no probe, no previous radiographs for comparison. Every one of those makes the unaided dentist worse than they are in real life, which makes the AI’s uplift look bigger than it is.
Who the readers were. Specialist dental radiologists and postgraduate researchers are not GDPs, and neither group behaves like a foundation dentist six weeks into practice.
The best UK-relevant example of a study that got the design right is the ADEPT trial, published in the British Dental Journal in 2021, which tested AssistDent from Manchester Imaging on enamel-only proximal caries in bitewings. Twenty-three dentists read cases with and without the software. Sensitivity rose from roughly 44% to roughly 76%. Specificity fell from roughly 98% to roughly 91%.
Note that the paper published the harm alongside the benefit. Most packs publish only the first number. Here is why the second one deserves your attention, worked through at a prevalence you might plausibly see:
Assume 1,000 proximal surfaces screened, 5% with true enamel-only caries
(50 diseased, 950 sound)
UNAIDED sens 44%, spec 98%
true positives 50 x 0.44 = 22
false positives 950 x 0.02 = 19
PPV = 22 / (22 + 19) = 54% -> roughly 1 in 2 flags is real
AI-AIDED sens 76%, spec 91%
true positives 50 x 0.76 = 38
false positives 950 x 0.09 = 86
PPV = 38 / (38 + 86) = 31% -> roughly 1 in 3 flags is real
Net: 16 more lesions found, 67 more false flags, per 1,000 surfaces
Sixteen extra early lesions caught per thousand surfaces is a genuinely good clinical outcome, particularly for preventive management. Sixty-seven extra false flags is a genuinely real cost in chair time, patient anxiety, and the temptation to intervene on sound enamel. Whether that trade is worth it in your practice depends on your case mix and your prevention pathway, which is a clinical judgement you can only make if the pack gives you both numbers. Run this arithmetic on any pack you are handed. It takes four minutes and it is the single most revealing thing you can do with a vendor’s figures.
Weak spot three: the comparison group
This is the one buyers skip, and it is where the money goes wrong.
Every uplift claim is a comparison against something. Look hard at what. “Improved detection versus unaided dentists” invites the question: which dentists, working how? If the comparator cohort were general dentists in a US corporate setting with a 30-minute exam slot, a full-mouth series at recall, and a payment model that rewards restorative intervention, that cohort’s baseline behaviour has almost nothing to do with a UK NHS UDA practice working to the College of General Dentistry’s selection criteria, taking bitewings at intervals set by risk, and operating under a contract that makes over-detection financially pointless and clinically indefensible.
The same problem shows up in triage tools with different clothing. A patient triage product will report agreement with a clinician on a set of presenting complaints. Ask what the clinician saw: free-text symptom descriptions in a clean dataset, or the actual mess of an out-of-hours message at 11pm. Then ask the only safety question that matters for triage, which is not accuracy but undertriage rate. How often did it route something that needed urgent care into routine? What was the worst case in the dataset, and what happened to it? A tool with 92% agreement and a 4% undertriage rate on pain-with-swelling presentations is a liability, however good the top-line looks.
Front-desk and admin tools get the loosest treatment of all, because they usually fall outside medical device regulation entirely and so face no evidential bar. A claim of “saves 45 minutes per day” needs three follow-ups: how was the baseline measured, over how many practices, and for how long. Pilots of three practices over two weeks measure novelty, not workflow. If there is voice charting or ambient note generation involved, ask for the word error rate on UK accents and on dental terminology, separately, because the general-purpose speech figures vendors quote are from datasets with neither.
The questions to send before the second meeting
Put these in an email. The quality of the reply is itself evidence, and a vendor with a real evidence base answers within a week.
| What the pack claims | What it actually establishes | What to ask |
|---|---|---|
| AUC 0.9x on test set | The model fits its own data distribution | Site-level breakdown of training vs test; was the test set external? |
| FDA 510(k) clearance | Substantial equivalence to a predicate device | Which predicate, which indication, which image types and ages? |
| CE / UKCA marked | Conformity assessment passed | Risk class, notified body, and MHRA registration number to check on the public database |
| Dentists found X% more caries | Aided readers flagged more, under study conditions | Specificity change, reader seniority, prevalence in the test set, washout design |
| Used by N,000 clinicians | Sales traction | How many UK sites, and can we speak to two on the NHS without you present? |
| Saves N minutes per day | Something was measured once | Baseline method, number of sites, duration, and what happened in month three |
Two administrative items belong on the same email and are easy wins in an NHS or mixed setting. Ask for their completed NHS Digital Technology Assessment Criteria submission and their Data Security and Protection Toolkit status. Ask where images are processed and stored, because a US-hosted inference endpoint changes your data protection impact assessment. Also confirm the regulatory route into Great Britain: CE marked devices remain acceptable here under the MHRA’s extended timelines rather than indefinitely, so a vendor with no UKCA plan is a vendor with a cliff edge in their roadmap.
Once the answers come back, the decision stops being about the algorithm and becomes a workflow question: who reviews the AI’s flags, what happens when clinician and software disagree, how it is recorded in the notes, and what you tell the patient. That sits alongside integration, training and contract terms in the wider buying decision, which is covered in Choosing and Integrating Tools.
One last thing worth saying plainly. When a flagged lesion turns out to be sound enamel, or an undertriaged patient presents with a spreading infection, the vendor’s evidence pack is not the document that gets examined. Your notes are, and your name is on them. Read the pack like the person who will be asked to defend the decision, because that is who you are.