PofoliaShared via Pofolia

Diagnostics· 2026Q2

Problem Formulation Outweighs Loss-Function Choice in Deep-Learning Cervical Vertebral Maturation Staging: A Multi-Seed Study on an Imbalanced Public Benchmark

Nazlı Tokatli

Short summary

Reformulating cervical vertebral maturation (CVM) staging from six stages to three clinically meaningful groups roughly doubled macro-F1 scores (to ~0.50) on an imbalanced dataset, with no significant difference between standard or specialized loss functions.

AI-generated from the title and abstract; the full text is not read.

Key points

  • Reformulating CVM staging from six stages to three clinically meaningful groups (pre-peak, peak, post-peak) significantly improved performance on an imbalanced dataset.
  • In the three-group CVM staging task, no loss function (unweighted cross-entropy, weighted cross-entropy, focal loss, ordinal loss) significantly outperformed others across eight seeds.
  • The unweighted baseline achieved a macro-F1 of 0.507 ± 0.045, comparable to the best specialized ordinal loss (0.505 ± 0.044).
  • Extreme imbalance in the six-stage task led to poor minority-stage recognition (macro-F1 ≈ 0.22–0.25).

AI-generated from the title and abstract; the full text is not read.

Abstract

(1) Background: Cervical vertebral maturation (CVM) staging on lateral cephalograms informs the timing of orthodontic treatment, and automated deep-learning approaches are widely studied. A recurring but under-examined obstacle is that public CVM datasets are severely imbalanced across the six stages. Using a public dataset and a rigorous eight-seed protocol, we quantify how strongly imbalance constrains performance, whether common imbalance-handling and ordinal-aware loss functions overcome it, and whether a clinically motivated three-group reformulation is more tractable than the native six-stage task. (2) Methods: Using the public Aariz dataset (1000 lateral cephalograms with expert CVM-stage labels; 700/150/150 train/validation/test, one radiograph per patient), we trained an ImageNet-pretrained ResNet-18 under four strategies: unweighted cross-entropy (baseline), class-weighted cross-entropy with balanced sampling, focal loss, and a rank-consistent ordinal (CORAL) head. Each configuration was trained with eight random seeds. Both tasks were evaluated over this multi-seed protocol, and sensitivity analyses examined the ordinal model’s sampling scheme and the focal-loss focusing parameter. Performance was assessed by accuracy, macro-F1, mean absolute error in stages (MAE), and quadratic weighted kappa (QWK), and compared across strategies with the Kruskal–Wallis test; secondary pairwise contrasts were Holm-corrected, and Wilson binomial confidence intervals (accuracy), seed-level confidence intervals (macro-F1), and analytically derived majority and frequency-weighted random classifiers were added as trivial reference baselines. Both the native six-stage task and a three-group scheme (pre-peak = CS1-3, peak = CS4, post-peak = CS5-6) were evaluated. (3) Results: On the six-stage task the extreme imbalance (CS1 n = 18 vs. CS5 n = 311 in training) produced poor minority-stage recognition (macro-F1 ≈ 0.22–0.25), with the rarest stages essentially unclassified even after imbalance handling. Reframing into three clinically meaningful groups roughly doubled macro-F1 (to ≈0.47–0.51) across all strategies. Critically, in the three-group setting no loss function significantly outperformed the others: across eight seeds the Kruskal–Wallis test was non-significant for every metric (macro-F1 p = 0.09, QWK p = 0.05, MAE p = 0.07, accuracy p = 0.41), and the unweighted baseline matched the best specialized loss on macro-F1 (0.507 ± 0.045 vs. ordinal 0.505 ± 0.044). In secondary pairwise comparisons the ordinal loss nominally exceeded the weighted and focal losses (uncorrected p = 0.04 and p = 0.01), but after Holm correction only the comparison with focal loss remained significant (adjusted p = 0.03), and no pairwise difference survived a conservative six-comparison Bonferroni bound; focal loss was the least stable. All strategies clearly exceeded trivial baselines on macro-F1 (majority classifier 0.259; frequency-weighted random ≈0.333), although the majority classifier’s accuracy (0.633) exceeded that of every trained model. (4) Conclusions: On this imbalanced public dataset, the dominant factor associated with CVM classification performance was the problem formulation rather than the loss function: a clinically grounded three-group scheme was far more tractable than six-stage classification, whereas specialized imbalance and ordinal losses did not reliably outperform a simple baseline. These findings, established under a fixed backbone, training budget, and a single public dataset, caution against assuming that loss-level fixes resolve CVM imbalance, and highlight data scale and problem formulation as the primary levers. Multi-dataset external validation is a necessary next step.

The authors' abstract, as published at the source. Diagnostics, 2026 · DOI ↗

TakeawaysPremium
Ask the paperFree account

Continue with a free account

Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.

Continue free on the web

Sign in with Google or Apple; no card needed. You come back to this paper.

On your phone:

Field: Orthodontics

OrthodonticsDentistry