AI Radiograph Reading
A practice principal in Leeds told me she’d cancelled her AI radiograph subscription after four months. Not because it was wrong. Because every bitewing set came back with three or four enamel-only lesions flagged in orange, her two associates started restoring them, and her composite usage went up 40% on a UDA contract that paid her exactly the same either way.
That is the real problem with ai dental x-ray analysis in a UK setting. The detection question is largely settled. The questions that decide whether you make or lose money are about false-positive burden, contract mix, integration with whatever imaging software you’re stuck with, and whether your clinical governance file can survive a CQC inspection or a GDC complaint. This page is about those.
What the software actually does to a bitewing
Strip away the marketing and a caries-detection model is doing object detection on a greyscale image. It was trained on tens or hundreds of thousands of intraoral radiographs where dentists (usually two or three, adjudicated) drew boxes around lesions. The model outputs a set of bounding boxes or pixel masks, each with a class label and a confidence score between 0 and 1. A display threshold decides which of those you actually see.
That threshold is the single most important setting in the product, and most vendors bury it. At 0.35 you will see everything including noise. At 0.70 you will see fewer, more reliable flags and miss borderline enamel lesions. The sensitivity and specificity figures in the vendor’s brochure were measured at one specific threshold, usually the one that makes the sensitivity number look best.
Here is roughly what a report looks like when you export it rather than viewing the overlay:
Patient: 100482 Image: BW-RIGHT Captured: 2026-09-14 Model v4.2
FINDINGS confidence
UR6 distal caries, into dentine 0.91
UR5 mesial caries, enamel only 0.62
LR6 mesial caries, into dentine 0.88
LR7 distal calculus 0.71
UR6 occlusal existing restoration 0.97
LR6 furcation bone level 4.1 mm from CEJ 0.84
LR7 periapical no radiolucency detected --
QUALITY: contact overlap UR5/UR6 = 0.8 mm (proximal read degraded)
Note the last line. Good products tell you when the image geometry has defeated them. A bitewing with 0.8 mm of overlap at the 5/6 contact is not readable by a model or by you, and the correct response is to retake it with a holder, not to trust a 0.62.
The tools UK practices are actually buying
| Tool | Origin | Covers | Typical UK integration route | Notes |
|---|---|---|---|---|
| Pearl Second Opinion | US | Intraoral caries, calculus, periapical radiolucency, bone level, restorations | Direct plugins for several imaging platforms plus a watched-folder bridge | Broadest condition list; also sells retrospective archive scanning |
| Overjet | US | Caries, quantified bone loss in mm, calculus | Deep integration with US practice management, thinner in UK | Strongest on perio measurement and reporting |
| AssistDent | UK (Manchester Imaging) | Proximal caries on bitewings only, with an enamel-lesion focus | Lightweight desktop overlay | Narrow scope on purpose; UK-developed, UK support |
| Diagnocat | International | 2D plus CBCT segmentation, full radiological report | Cloud upload, DICOM in | The only one on this list you’d buy primarily for CBCT |
| VideaHealth | US | Caries, with paediatric-specific models | Cloud with imaging integrations | Markets heavily on reducing missed caries rates |
| dentalXrai Pro | Germany | Panoramic and intraoral findings | Tight with Dürr VistaSoft | Useful if you’re already a Dürr practice |
Pricing in the UK sits, in quotes I’ve seen, somewhere between £90 and £250 per surgery per month, sometimes with a per-site floor and a setup fee in the £500 to £1,500 range. Archive-scanning projects are quoted separately. Get the quote in writing with the surgery count, the contract length, and what happens to the price at renewal, because three-year lock-ins at introductory rates are common.
Accuracy claims, and the number the brochure won’t show you
The most-cited UK evidence is the ADEPT study (Devlin and colleagues, Journal of Dentistry, 2021), which tested AssistDent on enamel-only proximal caries. Dentists reading unaided had a sensitivity of 24.0%. With the AI overlay, sensitivity rose to 71.0%. Specificity fell from 95.6% to 85.0%.
Those two movements are the entire story of AI radiograph reading, and the second one gets almost no airtime. Run the arithmetic on a real patient.
An adult four-film bitewing set gives you roughly 32 to 40 assessable proximal surfaces. Suppose 8% of them have an enamel-only lesion, which is a reasonable figure for a moderate-risk adult on a two-yearly radiograph interval. Take 36 surfaces: about 3 diseased, 33 sound.
- True positives: 3 × 0.71 = 2.1
- False positives: 33 × 0.15 = 5.0
So five flags on sound surfaces for every two real lesions found. Positive predictive value is 2.1 / 7.1, about 30%. Seven out of ten orange boxes on that screen are pointing at healthy enamel.
| Prevalence per surface | Sensitivity 71% | Specificity 85% | PPV |
|---|---|---|---|
| 3% | 0.021 | 0.146 | 13% |
| 8% | 0.057 | 0.138 | 29% |
| 15% | 0.107 | 0.128 | 45% |
| 30% | 0.213 | 0.105 | 67% |
This is why the same tool feels brilliant in a high-caries community practice in Blackpool and feels like noise in a low-caries private list in Harrogate. The model didn’t change. Your prevalence did. We go through the bitewing evidence in much more detail, including how the reference standard was set and why histological validation studies give different numbers again, in How Accurate Is AI Caries Detection on Bitewings?.
The false-positive tax
Five spurious flags per patient sounds tolerable until you multiply it by throughput. An associate seeing 1,800 patients a year, taking bitewings on half of them, is looking at roughly 900 sets and around 4,500 false flags annually.
Each one costs something. Thirty seconds of looking, sometimes a second opinion from a colleague, sometimes a conversation with the patient that starts “the computer has picked up something”. If the average is 40 seconds of clinical and administrative time, that’s 50 hours a year per associate. On a chair costing you £120 an hour to run, £6,000 of time to handle findings that mostly shouldn’t change management.
The clinical risk is worse than the time cost. Enamel-only proximal lesions should not be restored. Delivering Better Oral Health (4th edition) and the College of General Dentistry selection criteria both point to non-operative management: 22,600 ppm fluoride varnish twice yearly, dietary advice, interdental cleaning, radiographic monitoring at an appropriate interval. An AI tool that reliably surfaces E1 and E2 lesions has handed you a prevention workload, not a restorative one, and if your associates convert those flags into preparations you have an overtreatment problem with a documented audit trail showing exactly why each cavity was cut.
Set the display threshold high for your first quarter. 0.70 or above. You can always lower it once your team has calibrated.
Where the NHS contract breaks the ROI case
Vendor ROI calculators are built on US fee-per-item economics. They assume every additional detected lesion becomes billable treatment. Under an English NHS contract, a Band 2 course of treatment is 3 UDAs whether you place one filling or six. The national minimum UDA value has been £28.00 since April 2024, with most contracts sitting somewhere between £28 and £40.
Work it through for a mixed practice. Call it three surgeries, 6,200 patients, 72% NHS by course-of-treatment volume, UDA value £32.
NHS side. AI surfaces, say, 1.3 additional restorable dentinal lesions per 100 examinations that would otherwise have been missed this cycle. Across roughly 4,400 NHS courses of treatment, that’s 57 extra restorations. Additional UDA income: close to zero, because nearly all of them fall inside a Band 2 that was already being claimed. Additional chair time at 20 minutes each: 19 hours. On NHS work, better detection is a cost.
Private and plan side. The same detection rate across 1,700 private courses of treatment gives about 22 additional restorations at £185, so roughly £4,070 of fee income, plus whatever flows from bone-loss findings into hygiene and periodontal therapy.
Cost. Three surgeries at £150 per month is £5,400 a year, plus a £900 setup.
The first year is underwater on those assumptions. The case gets built elsewhere: in periodontal staging, in retrospective archive scanning that surfaces untreated disease in your existing recall base, and in the defensibility of your records. Be honest with yourself about which of those you’re actually buying, because the detection story alone does not pay for itself on a heavily NHS list.
Bone loss is the quietly useful bit
Since the British Society of Periodontology’s implementation of the 2017 classification, staging depends on radiographic bone loss at the worst site, expressed as a percentage of root length. Stage I is up to 15%, Stage II is 15% to 33%, Stage III extends into the middle third and beyond, Stage IV adds the tooth-loss and functional criteria.
Measuring that by hand means identifying the CEJ, the most apical extent of bone, and the root apex, then doing the division. On a full-mouth periapical series that is a genuinely unpleasant ten minutes, which is why in practice a lot of staging gets eyeballed. Overjet and Pearl both output CEJ-to-bone distances in millimetres per site, and once you have those numbers the percentage falls out:
bone loss % = (CEJ-to-bone distance) / (CEJ-to-apex distance) x 100
LR6 mesial: 4.1 mm / 13.8 mm = 30% -> Stage II
UR7 distal: 6.9 mm / 14.2 mm = 49% -> Stage III
This is the use case where the AI is doing a measurement task rather than a judgement task, and measurement is what machines are good at. It also produces something you can show a patient and put in a referral letter. If you’re evaluating tools and your list skews towards periodontal disease rather than caries, weight the perio module heavily in your trial.
Integration is what will actually stop you
Nobody’s pilot fails because the model is bad. They fail because the images can’t get to it.
Find out, before any demo, what your imaging chain is. Sidexis 4 with Schick sensors, Romexis with Planmeca, CS Imaging with Carestream, VistaSoft or DBSWIN with Dürr PSP plates, Apteryx XVWeb, Dexis. Then ask the vendor a precise question: does the integration read from the imaging database directly, does it hook the acquisition event, or does it watch an export folder?
Watched folders work, but they mean someone has to export. Within about three weeks of go-live, exporting stops happening on busy days, and you are paying for a tool nobody uses. Acquisition hooks are what you want: image captured, AI runs, overlay appears without a human decision.
Ask also where the inference happens. Cloud processing means patient radiographs leave your building, which brings in the data protection work below. On-premise inference needs a machine with a GPU, and if the vendor says “any modern PC will do”, ask for the actual minimum spec and the per-image latency on it. Anything over about 15 seconds per bitewing will not survive contact with a real surgery.
The paperwork you need before go-live
These tools are medical devices. Not “arguably” or “probably”: software intended to detect disease from an image is a device, and the detection features you’re buying sit under UK MDR 2002 as amended.
Get these from the vendor in writing:
- UKCA mark, or a CE mark you can rely on. MHRA currently recognises CE marking, with different end dates depending on which route the certificate came through. Ask for the certificate number and the notified body, and check the end date of recognition against your contract length.
- The intended purpose statement, verbatim. It will almost certainly say the software is an aid to the dentist and that the dentist retains responsibility for diagnosis. That sentence is what you put in your governance file, and it is what your practice protocol must reflect.
- A clinical evaluation summary. Which populations, which image types, which reference standard. If it was validated on US sensor images and you shoot PSP plates, say so out loud and factor it into your own validation.
- A DPIA, and a processor agreement under Article 28. If images go to a cloud service outside the UK, you need a transfer mechanism (IDTA or the UK Addendum) and a transfer risk assessment. The practice is the controller. That responsibility does not move.
- DTAC and your DSPT position, if you are dealing with any NHS commissioner-facing procurement or integrating with NHS systems.
- A written answer from your indemnifier. Dental Protection, DDU and MDDUS will all engage on this. The question to ask is specific: does using this tool as a concurrent second read, within its stated intended purpose, affect cover.
One more thing on IR(ME)R 2017. AI output is not a justification for an exposure. Every radiograph still needs justification by the practitioner before it is taken, on clinical grounds, and “the AI wants a comparison image” will not hold up. Nor does an AI tool remove the need for your radiograph quality audit against the current selection criteria targets.
A 90-day evaluation that produces evidence
Most trials are two weeks of enthusiasm followed by a purchase decision made on vibes. Do this instead.
Weeks 1 to 2, blinded baseline. Pull 40 anonymised bitewing sets from your archive, weighted to your actual case mix. Have two clinicians independently score every proximal surface without AI. Record disagreements. This is your unaided baseline, and it is also a genuinely uncomfortable exercise, because inter-examiner agreement in general practice is usually worse than people expect.
Weeks 3 to 4, AI on the same 40 sets. Same clinicians, at least a fortnight later so they’ve forgotten. Record what the AI added, what it missed that the clinicians caught, and how many flags were rejected. You now have your own local sensitivity gain and your own false-flag rate, on your images, from your sensors.
Weeks 5 to 10, live with logging. Threshold at 0.70. One line per patient in a spreadsheet: flags shown, flags accepted, management change (none, prevention, restoration, referral). Ten minutes a day between the whole team.
Weeks 11 to 12, the numbers. Cost per useful finding. Chair-time delta. Split by NHS and private. Then decide.
Vendor questions, sent as one email:
1. Exact intended purpose statement as it appears on the label.
2. UKCA/CE certificate number, notified body, recognition end date.
3. Sensitivity and specificity by condition, with the threshold each
figure was measured at, and the reference standard used.
4. Sensor and plate types in the training data. Any PSP?
5. Integration method with [our imaging software, version]:
acquisition hook, database read, or watched folder?
6. Inference location. If cloud: country, sub-processors, retention,
and do you train on our images? (Get the answer in the contract.)
7. Per-image latency on our hardware spec.
8. Can we set and lock the confidence threshold practice-wide?
9. What happens to overlays and reports if we cancel?
10. Notice period, price at renewal, and cost per extra surgery.
Talking to patients without turning it into a sales tool
Showing someone an annotated radiograph is persuasive. That is exactly the problem. A patient looking at orange boxes on their own teeth will agree to almost anything, and GDC Standard 3.1 requires valid consent based on adequate information, which includes the information that most of those boxes represent early enamel changes best managed with fluoride varnish and better interdental cleaning.
Give your team a script. Something like: “The software has highlighted some early changes in the enamel here. Nothing that needs drilling. What it means is that I want to put a high-fluoride varnish on today, get you cleaning between these two teeth properly, and take another picture in eighteen months to check it’s stable.” That sentence is clinically correct, it’s defensible, and it still demonstrates value to the patient.
Where AI-annotated images genuinely earn their keep is in the conversations you currently lose: the patient who won’t accept that a dentinal lesion under an old amalgam needs dealing with, and the periodontal case where a bone-loss percentage per site makes the referral make sense.
The Leeds practice went back to AI nine months later, incidentally. Same vendor, threshold locked at 0.75 by the principal rather than left on the default, a written protocol saying enamel-only findings go to the prevention pathway and not the restorative one, and the perio module switched on for the hygienists. Their composite usage settled back to baseline and their periodontal therapy revenue went up 18%.
In this section
The supporting pages under this subject.