PofoliaShared via Pofolia

Diagnostics· 2026Q2· Review

Reference Standards in Artificial Intelligence Detection of Apical Periodontitis: A Narrative Review of What Reported Accuracy Measures

Omar Rifat Alkhattab, Abdullah Bokhary

Short summary

AI detection of apical periodontitis often misinterprets accuracy, with reported sensitivities (61.0%) and negative predictive values (41.6%) based on periapical radiography not reflecting true disease detection, which is significantly lower (16-27% without parallax, 38% with).

AI-generated from the title and abstract; the full text is not read.

Key points

  • Reported AI accuracy for apical periodontitis often uses human readings of the same image as a reference, not actual disease presence.
  • Periapical radiography detects only 16-38% of confirmed lesions compared to histopathology.
  • AI models achieve pooled digital sensitivities of 61.0% and negative predictive values of 41.6% based on these less stringent references.
  • Detection accuracy is influenced by lesion size, with smaller lesions (<1-3 mm) being invisible in incisors and molars, respectively.

AI-generated from the title and abstract; the full text is not read.

Abstract

The accuracy reported for detecting apical periodontitis using artificial intelligence (AI) is generally interpreted as accuracy against the disease; however, this is usually not the case. Compared with histopathology, single-view periapical radiography detects 16 to 27% of confirmed lesions and 38% with parallax, with a pooled digital sensitivity of 61.0% and a negative predictive value of 41.6%. Two blinded specialists reading 1717 radiographs disagreed on 22% of teeth. Detection depended on lesion size and site: simulated cancellous-bone lesions were invisible below 1 mm in incisors and 3 mm in molars. Of the 44 primary studies identified, 34 (77%) used a label that was a reading of the index image or was unascertainable; eight (18%) used an independent reference—six (14%) suited to contemporaneous detection and two to later treatment outcomes. No cone-beam computed tomography (CBCT) input study had a reference independent of its volume. One commercial platform reported 92.3% sensitivity against same-image consensus and 47.9% against blinded CBCT. The reported accuracy is thus in agreement with a noisy rater; recognising this changes what meta-analyses pool. A model that reproduces human labels would have a lower sensitivity against disease than reported; for real-world models, the direction depends on error dependence that the study design cannot identify.

The authors' abstract, as published at the source. Diagnostics, 2026 · DOI ↗

TakeawaysPremium
Ask the paperFree account

Continue with a free account

Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.

Continue free on the web

Sign in with Google or Apple; no card needed. You come back to this paper.

On your phone:

Field: Oral Surgery

Oral SurgeryDentistry