PofoliaShared via Pofolia

International Endodontic Journal· 2026Q1

Open‐Weight Deep Learning for Periapical Index Scoring: Training and Comparison of Convolutional Neural Network Models Based on Three Distinct Architectures and an Ensemble Approach

Gerald Torgersen, Karianne Winsnes, Trude Handal, My Tien Diep et al.

Short summary

An ensemble of three CNN models achieved the highest performance (QWK=0.79) for Periapical Index (PAI) scoring on dental radiographs, though not statistically significantly better than individual models.

AI-generated from the title and abstract; the full text is not read.

Abstract

ABSTRACT Aim This study aimed to (i) train and compare three convolutional neural network (CNN) architectures for Periapical Index (PAI) scoring on intraoral radiographs, and (ii) assess whether an ensemble of these models improves performance compared with any single architecture. Methods A dataset of 17,549 apex‐centred image clips extracted from intraoral radiographs was annotated by a calibrated observer (KW) with PAI scores (1–5) and used to train three CNN architectures (ResNet50, EfficientNet‐B3 and ConvNeXt‐Tiny) for PAI scoring. The image clips were split 80/20 into training and validation sets, and model selection was based on validation performance of quadratic weighted kappa (QWK). The three CNN models were combined into an equal‐weight soft‐voting ensemble model, and all four models were evaluated on an independent balanced test set of 200 images (40 per PAI score). QWK was used as the primary outcome metric, and bootstrapping was used to estimate the confidence intervals (CIs). The PAI scale was dichotomized into healthy (PAI 1–2) and diseased (PAI 3–5), and accuracy, sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV) and F1 score were calculated. Results The ensemble model achieved the highest performance with a QWK of 0.79 on the test set, followed by EfficientNet‐B3 (QWK = 0.77), ResNet50 (QWK = 0.76) and ConvNeXt‐Tiny (QWK = 0.70), with overlapping confidence intervals (CIs). For the binary PAI classification, the performance of the three CNN models and the ensemble model was comparable. The ensemble model achieved the following results: accuracy 86%, sensitivity 83%, specificity 90%, PPV 93%, NPV 78% and F1 score 0.88. Conclusion Optimization and evaluation of three CNN models and an ensemble model for PAI scoring showed comparable performance on both the five‐class PAI scale and the binary PAI scale on an independent test set. The ensemble model did not show statistically significant improvement over the three individual CNN architectures. The AI models trained and tested in this study are made open‐source under the open‐source MIT licence.

The authors' abstract, as published at the source. International Endodontic Journal, 2026 · DOI ↗

TakeawaysIn the app
Key pointsIn the app
Ask the paperIn the app

The rest is in the Pofolia app

Takeaways, key points and questions to the paper; new summaries every day for your field. Free.

Sign in on the web to open

Field: Oral Surgery

Oral SurgeryDentistry