AI Dent
§2 Section 2 of 6 2,492 words · 11 min

Choosing and Integrating Tools

Most dental AI software comparison exercises go wrong in the first ten minutes, because someone builds a spreadsheet with “accuracy” as a column. Three vendors write 95% in it. The spreadsheet now contains no information at all.

What follows is the comparison framework I’d use for a UK practice buying into three genuinely different product categories: radiograph interpretation, patient triage, and front-desk admin. They have different regulatory status, different failure modes, different pricing logic and wildly different integration difficulty. Treating them as one purchase is the most expensive mistake available to you.

The three categories are not comparable, so stop comparing them

Radiograph AI is a medical device. In Great Britain that means UKCA marking, or a CE mark you are still permitted to rely on (devices certified under EU MDR are accepted in GB until 30 June 2030; older MDD and AIMDD certificates until 30 June 2028, or certificate expiry, whichever is sooner). Northern Ireland sits under EU rules. Caries detection and bone-level measurement products typically land at Class IIa, which means a notified body or approved body has looked at the clinical evidence. Ask for the certificate number and the exact intended-use statement, not a brochure.

Front-desk admin tools are not medical devices and shouldn’t pretend to be. An AI that answers the phone, books a hygiene appointment and texts a confirmation carries commercial and data protection risk, not clinical risk. Procurement should be correspondingly lighter: you need a solid Article 28 processor agreement, a transfer mechanism if the vendor hosts outside the UK, and a way to get your data out. You do not need a notified body certificate.

Triage is the awkward middle. Sorting an inbound “my tooth hurts” message into urgent, routine, or not-a-dental-problem is a clinical decision when it determines whether someone is seen today. Vendors will tell you their triage tool is “not diagnostic” and therefore out of scope. Sometimes true, sometimes a convenient reading of the intended purpose. The test to apply: if the software’s output changes when a patient is seen, and no registered clinician reviews that output before it takes effect, you have built an unregulated diagnostic pathway and you own it.

Radiograph AI: sensitivity and specificity, never “accuracy”

The best publicly available UK number on this comes from the ADEPT study (Devlin and colleagues, Journal of Dentistry, 2021), which tested Manchester Imaging’s AssistDent on enamel-only proximal caries in bitewings. Unaided dentists detected roughly 44% of enamel-only lesions. With the software, that rose to about 60%. Specificity fell from roughly 99% to 87%.

That trade is the whole story, and it is worth doing the arithmetic on your own volumes. Take a three-surgery mixed practice taking 1,500 bitewing pairs a year. Score twenty proximal surfaces per pair and you are looking at 30,000 surfaces annually. If 10% carry an enamel-only lesion:

  • Unaided: about 1,330 of 3,000 lesions detected, and around 378 false flags on sound surfaces.
  • AI-assisted: about 1,810 lesions detected (+480), and around 3,456 false flags (+3,078).

So roughly one extra genuine early lesion for every six extra false prompts. Whether that’s a good deal depends entirely on what your team does with a flag. If a flag triggers a fissure sealant conversation and a six-month review, it’s cheap. If a flag triggers a restoration, you have industrialised overtreatment and your indemnifier will eventually ask you about it.

Different products make different claims and the claims matter more than the brand. Pearl’s Second Opinion covers caries, periapical radiolucency, calculus, restorative margin discrepancies and bone level measurement across 2D imaging. Overjet’s strength is quantified bone loss in millimetres, which makes perio staging auditable in a way clinician eyeballing isn’t. VideaHealth’s caries and perio modules are built around per-tooth confidence scoring. Diagnocat’s value is CBCT: automated segmentation, per-tooth findings, third molar and canal reporting, on a credit-per-study basis. dentalXrai Pro out of Berlin ships tightly bound to Dürr’s VistaSoft, which is either a feature or a cage depending on your imaging estate.

A real output looks less like a diagnosis and more like a worklist:

Patient 10482  |  Bitewing L  |  2026-09-14  |  model v4.2
UR6 distal    caries, dentine        conf 0.91   [agree]
UR5 mesial    caries, enamel         conf 0.62   [agree] [reject]
UR4 distal    caries, enamel         conf 0.41   below display threshold
UR6           restoration margin     conf 0.77   [agree] [reject]
LR7-LR6       bone level 4.1 mm      CEJ-to-crest, mesial LR6
Clinical evaluation: NOT RECORDED  <-- IR(ME)R 2017 outstanding

That last line is the one to design your workflow around. Under IR(ME)R 2017 every exposure needs a recorded clinical evaluation by an entitled practitioner. An AI report is not a clinical evaluation and no vendor will claim otherwise. If your notes end up containing only the machine output, you have fewer records than you had before you bought the software. Set the display threshold deliberately too: most products let you tune it, and running at the vendor default because nobody knew it was configurable is common.

What a pilot has to produce to be worth running

Six weeks, one clinician, one metric that changes a decision. Anything larger stalls.

QuestionHow to measure itWhat “pass” looks like
Does it find things we missed?Retrospective run over 200 archived bitewings with known 12-month outcomes≥15 lesions flagged that were later treated
How noisy is it?Count flags rejected at the chair per 100 images<25 rejections per 100 images
Does it slow us down?Stopwatch on 30 consecutive radiograph reviews, with and without<20 seconds added per image
Do patients respond differently?Treatment plan acceptance rate on flagged early lesionsMeasurable shift, either direction
Does it survive the PMS?Count images that failed to reach the AI unattendedZero manual exports in week 6

The retrospective run is the part vendors resist and the part you should insist on, because it’s the only test where you already know the answer. Ask for a time-limited licence against exported DICOM, and make it a condition of the trial. If the product can’t ingest your archive, that’s a finding.

Pricing models change which product is cheapest

Two dominant shapes. Per-site subscription: a flat monthly figure per surgery or per location, unlimited or very high image allowance. Credit-based: you buy analyses, 2D studies cost a little, CBCT a lot more.

Run both against your actual radiograph count. A practice at three surgeries taking 1,500 bitewing pairs, 700 periapicals and 400 panoramics does roughly 2,600 studies a year:

Per-site subscription  £420/month x 12          = £5,040/yr  -> £1.94 per study
Credit model           2,600 x £2.80            = £7,280/yr  -> £2.80 per study
Break-even                £5,040 / £2.80        = 1,800 studies/yr

Below 1,800 studies the credit model wins. Above it the subscription does, and the gap widens every year as your imaging volume grows. A single-surgery squat practice taking 600 studies should be nowhere near a flat per-site licence. Quotes in the UK market for 2D radiograph AI cluster somewhere between £250 and £600 per site per month depending on volume, module count and contract length, with CBCT reporting priced separately because the compute cost is real.

Watch for three things in the paper. Per-surgery pricing that counts chairs you don’t use for imaging. Automatic uplifts pegged to something other than CPI. And minimum terms of 36 months on a product category where the model you’re buying will be replaced twice in that window.

Triage: build the red flags as rules, not as a model

The genuinely useful triage function in a UK practice is not diagnosis, it’s sorting: which of the 40 messages in the inbox needs a same-day emergency slot, which needs a routine appointment, which is a prescription request, and which is asking about parking. Classification at that level is reliable and saves a receptionist an hour a day.

Hard-code the escalations. Extra-oral swelling, difficulty swallowing or breathing, swelling closing the eye, trismus, temperature above 38°C, and any post-extraction bleeding that hasn’t stopped with pressure should route to a human immediately by keyword and structured question, never by model inference. A false negative on Ludwig’s angina is not a product defect you write up in a quarterly review.

Triage output also has to map onto NHS mechanics or it creates work rather than removing it. Urgent treatment is 1.2 UDAs. A Band 1 course is 1, Band 2a is 3, Band 2b is 5, Band 2c is 7, Band 3 is 12. If your triage tool books an “emergency” slot for something that turns into a Band 2c molar endodontic course, the scheduling assumption behind that slot was wrong and the clinician absorbs it. Ask the vendor to show you how the triage category maps to appointment type and duration in your diary, with your book, before you buy. UK teledentistry products like Toothfairy handle patient-facing triage reasonably well; practice-side enquiry routing is more often served by CRM tools such as DenGro, where the “AI” is a thin layer over rules that you can actually inspect.

Front-desk admin has the best return and the lowest clinical risk

This is where the money is, and it’s routinely deprioritised because it isn’t interesting.

Take the same three-surgery practice: 180 inbound calls a week, 22% unanswered at first attempt, concentrated 08:30 to 09:30 and just after lunch. That’s around 40 missed calls weekly. Roughly a third never call back, so 13 or 14 conversations a week simply evaporate. Suppose two of those were new private patient enquiries at an average first-year value of £350 (exam, two hygiene visits, one direct restoration). That’s £700 a week of enquiry value going to the practice down the road, £36,000 a year on paper.

Don’t budget the £36,000. An AI voice agent answering those calls will convert some fraction: assume 40%, so about £14,500 a year recovered. Voice products price per minute, typically £0.12 to £0.25, so 40 calls a week averaging three minutes is roughly £1,200 a year of usage plus a platform fee. The ratio holds up even if you halve the conversion assumption. Arini, Peerlogic and Weave all play in this space, all US-origin, which makes the data transfer question below non-optional.

Failure to attend is the second lever. At 240 appointments a week and a 9% FTA rate you lose about 22 slots weekly. Better reminder sequencing plus automated waitlist backfill routinely gets that to 6.5%, recovering six slots a week. At an average contribution of £62 per slot, that’s £19,000 a year of capacity you have already paid staff to be present for.

FP17 validation is the unglamorous third. A practice submitting 6,800 FP17s a year at a 1.8% rejection rate deals with 122 rejections; at twelve minutes each to investigate and resubmit, that’s 24 hours of someone’s year. A rules engine that catches band mismatches, missing exemption evidence and duplicate courses before submission pays for itself in a quarter and carries no clinical risk whatsoever.

Integration decides whether any of this works

The model quality is rarely the constraint. Getting images and patient records in and out of SOE Exact, Carestream R4+ or Dentally is the constraint, and the three systems behave nothing alike. Dentally is cloud-native with a documented REST API and OAuth 2.0. R4+ sits on Microsoft SQL Server with imaging handled through CS Imaging. Exact runs on a Sybase SQL Anywhere backend, and integrations that go around the official partner programme by reading the database directly tend to break on the next version upgrade. Full detail on all three, including what the official routes actually permit, is in Integrating AI With SOE, R4 and Dentally.

The practical rule: get the AI vendor and your PMS vendor on the same call before you sign anything. Not an email. A call, with your practice manager on it, where someone says out loud how images will reach the model and how findings will get back into the patient record.

Run a smoke test in week one of any pilot:

$ dcmsend -v ai-gateway.local 11112 +sd +r ./test-bitewings/
I: Requesting Association
I: Association Accepted (Max Send PDV: 16372)
I: Sending file: ./test-bitewings/IM0001.dcm
I: Converting transfer syntax: JPEG Lossless -> Little Endian Explicit
I: Received Store Response (Success)
...
I: Sent 24 of 24 objects

$ curl -s https://api.vendor.example/v1/studies?since=2026-09-29 | jq '.[].status'
"analysed"   x22
"rejected: missing PatientID"   x2

Those two rejections are the whole project in miniature. Images exported from R4 without a populated PatientID will not reconcile back to a patient, and the fix is a mapping table someone has to own. Better to find that on day three than in month four.

Contract terms worth arguing over

Data exit first. You want a written commitment that on termination you receive your radiographs and all generated findings in DICOM and a documented structured format, within a stated number of days, at no cost. Vendors who charge an export fee are telling you something about their retention strategy.

Then the transfer paperwork. A US-hosted vendor processing UK patient data needs an International Data Transfer Agreement or the UK Addendum to the EU SCCs, plus a transfer risk assessment you can show the ICO. Your DPIA under Article 35 is mandatory here, not advisory: automated processing of health data at scale hits the threshold squarely. If you hold an NHS contract you also complete the Data Security and Protection Toolkit annually, and your vendor list is part of that.

Ask about model versioning in writing. When the vendor pushes v4.3, are your thresholds preserved? Do you get release notes with performance deltas? Is there a rollback? A product that silently changes its sensitivity between Tuesday and Wednesday makes your clinical audit meaningless, and “continuous improvement” is not an answer to that question.

One more: check with Dental Protection, MDDUS or the DDU before you go live, and keep the reply. Their consistent position is that these tools are decision support and the registrant remains responsible, which is fine, but you want the specific confirmation on file rather than the general principle.

The mistake nearly everyone makes

Buying the radiograph AI first. It’s the most visible, the most fun to demo at a study day, and the one that generates the best conversation with patients. It is also the category with the heaviest regulatory load, the hardest integration, the longest payback and the most contested evidence base.

Start with the phones and the FP17s. Get a full quarter of clean data on missed calls, FTA rates and claim rejections, prove the operational discipline exists to act on what the software tells you, then take that same discipline to a radiograph pilot with a retrospective arm and a threshold you chose yourself. A practice that can’t reliably backfill a cancelled slot is not a practice that will handle 3,000 extra caries flags a year well.

In this section

The supporting pages under this subject.