Diagnostics· 2026Q2· Review
Reference Standards in Artificial Intelligence Detection of Apical Periodontitis: A Narrative Review of What Reported Accuracy Measures
- 0citations
- Q2SCImago
- 2026year
Short summary
AI detection of apical periodontitis often misinterprets accuracy, with reported sensitivities (61.0%) and negative predictive values (41.6%) based on periapical radiography not reflecting true disease detection, which is significantly lower (16-27% without parallax, 38% with).
AI-generated from the title and abstract; the full text is not read.
Key points
- Reported AI accuracy for apical periodontitis often uses human readings of the same image as a reference, not actual disease presence.
- Periapical radiography detects only 16-38% of confirmed lesions compared to histopathology.
- AI models achieve pooled digital sensitivities of 61.0% and negative predictive values of 41.6% based on these less stringent references.
- Detection accuracy is influenced by lesion size, with smaller lesions (<1-3 mm) being invisible in incisors and molars, respectively.
AI-generated from the title and abstract; the full text is not read.
Abstract
The accuracy reported for detecting apical periodontitis using artificial intelligence (AI) is generally interpreted as accuracy against the disease; however, this is usually not the case. Compared with histopathology, single-view periapical radiography detects 16 to 27% of confirmed lesions and 38% with parallax, with a pooled digital sensitivity of 61.0% and a negative predictive value of 41.6%. Two blinded specialists reading 1717 radiographs disagreed on 22% of teeth. Detection depended on lesion size and site: simulated cancellous-bone lesions were invisible below 1 mm in incisors and 3 mm in molars. Of the 44 primary studies identified, 34 (77%) used a label that was a reading of the index image or was unascertainable; eight (18%) used an independent reference—six (14%) suited to contemporaneous detection and two to later treatment outcomes. No cone-beam computed tomography (CBCT) input study had a reference independent of its volume. One commercial platform reported 92.3% sensitivity against same-image consensus and 47.9% against blinded CBCT. The reported accuracy is thus in agreement with a noisy rater; recognising this changes what meta-analyses pool. A model that reproduces human labels would have a lower sensitivity against disease than reported; for real-world models, the direction depends on error dependence that the study design cannot identify.
The authors' abstract, as published at the source. Diagnostics, 2026 · DOI ↗
Continue with a free account
Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.
Continue free on the webSign in with Google or Apple; no card needed. You come back to this paper.
On your phone:
Field: Oral Surgery
Oral SurgeryDentistry