PofoliaShared via Pofolia

BMC Oral Health· 2026Q1· Review

Artificial intelligence for radiographic assessment of peri-implant marginal bone loss: a systematic review and meta-analysis of model performance

Ying Wang, Yi Li, Feng Liu

Short summary

AI models achieved a pooled accuracy of 0.934 for radiographic assessment of peri-implant marginal bone loss, with pooled sensitivity of 0.930 and specificity of 0.961, based on a meta-analysis of five studies.

AI-generated from the title and abstract; the full text is not read.

Key points

  • AI models achieved a pooled accuracy of 0.934 (95% CI, 0.874–0.967) for radiographic assessment of peri-implant marginal bone loss.
  • Pooled sensitivity and specificity for AI models were 0.930 (95% CI, 0.876–0.961) and 0.961 (95% CI, 0.886–0.988), respectively.
  • Included studies used various imaging modalities (periapical, panoramic, CBCT) and AI architectures (CNNs, Faster R-CNN, U-Net, YOLO).
  • Most studies relied on internal validation, with limited external validation or prospective clinical workflow evaluation.
  • Substantial heterogeneity and inconsistent reporting limit the generalizability of current AI performance metrics.

AI-generated from the title and abstract; the full text is not read.

Abstract

Artificial intelligence has been applied to dental radiographic interpretation, including assessment of peri-implant marginal bone loss and peri-implantitis-related bone defects. Most available evidence has focused on lesion detection; quantitative measurement, severity grading, and clinical applicability have been less consistently examined. This systematic review and meta-analysis evaluated the performance of AI models for radiographic assessment of peri-implant marginal bone loss and related bone defects. PubMed/MEDLINE, Embase, Web of Science Core Collection, Scopus, Cochrane Library, and IEEE Xplore were searched from inception to 25 May 2026. Google Scholar, reference lists, and citation tracking were also used to identify additional records. English-language original studies were eligible if they used AI-based methods to assess peri-implant marginal bone loss, peri-implant bone-level changes, peri-implantitis-related radiographic bone defects, or severity grading on dental radiographic images. Study characteristics, imaging modality, AI architecture, target task, reference standard, validation strategy, model performance, and validation and clinical-translation features were extracted. Exploratory random-effects single-proportion meta-analyses were performed only when at least three studies reported sufficiently comparable internal performance metrics. Twelve original studies were included. Imaging modalities included periapical radiographs, panoramic radiographs or orthopantomographs, intraoral radiographs, and CBCT-derived images. By task category, detection and classification studies mainly used Faster R-CNN, YOLO-based models, AlexNet, and CNN-based classifiers; segmentation-related studies used U-Net or Mask R-CNN; measurement and keypoint-based studies used modified R-CNN, YOLOv8-pose, or image-processing pipelines; and one study used a multi-stage deep learning cascade workflow for longitudinal quantification. Exploratory pooling was performed only for selected internal performance metrics. Five studies contributed to the exploratory meta-analysis of model accuracy, with a pooled accuracy of 0.934 (95% CI, 0.874–0.967; I² = 96.4%). Three studies contributed to pooling of sensitivity/recall and specificity, yielding estimates of 0.930 (95% CI, 0.876–0.961; I² = 82.9%) and 0.961 (95% CI, 0.886–0.988; I² = 94.6%), respectively. Because heterogeneity was substantial, these estimates were interpreted as descriptive summaries of internal model performance rather than diagnostic accuracy for a single clinical endpoint. Precision, severity grading, measurement, keypoint-based, segmentation, and longitudinal quantification outcomes were summarized narratively because the tasks, classification schemes, and reported metrics were not sufficiently comparable for pooling. Most studies relied on internal validation, and none provided robust external validation or prospective clinical workflow evaluation. AI models generally showed promising internal validation performance for radiographic assessment of peri-implant marginal bone loss and peri-implantitis-related bone defects. However, the evidence remains limited by substantial heterogeneity, inconsistent reference standards, incomplete reporting of validation procedures, limited external validation, and a lack of prospective clinical workflow studies. Current AI models may serve as radiographic decision-support tools rather than stand-alone diagnostic systems for peri-implantitis.

The authors' abstract, as published at the source. BMC Oral Health, 2026 · DOI ↗

TakeawaysIn the app
Ask the paperIn the app

The rest is in the Pofolia app

Takeaways and questions to the paper; new summaries every day for your field. Free.

Sign in on the web to open

Field: Oral Surgery

Oral SurgeryDentistry