Bioengineering· 2026Q2
Interpretable and Calibrated Classification of Clinical Data Using Supervised Feature Binarization
- 0citations
- Q2SCImago
- 2026year
Short summary
A new framework uses supervised binarization to convert continuous clinical data into interpretable binary rules for Bernoulli Naïve Bayes models, achieving high AUC scores (up to 0.984) on benchmark datasets while ensuring reproducible predictions.
AI-generated from the title and abstract; the full text is not read.
Key points
- Supervised chi-square-guided binarization converts continuous clinical variables into binary indicators.
- Bernoulli Naïve Bayes (BNB) model is used for interpretable, rule-based clinical classification.
- Achieved AUC scores of 0.794 (Diabetes), 0.984 (Breast Cancer), and 0.918 (Heart Failure).
- Probabilistic reliability was assessed via cross-validated calibration, with post hoc beta calibration improving results.
- Model inference is reproducible from a printed reference table using basic arithmetic.
AI-generated from the title and abstract; the full text is not read.
Abstract
Black-box models limit the adoption of artificial intelligence in medicine because their predictions are difficult to interpret and reproduce. We present a statistically grounded framework for interpretable, rule-based clinical classification using the Bernoulli Naïve Bayes (BNB) model. Supervised chi-square-guided binarization converts continuous variables into binary indicators by selecting thresholds that maximize association with the clinical outcome within the training folds, which allows BNB to operate on continuous medical data without sacrificing transparency. On three benchmark datasets, Pima Indians Diabetes, Wisconsin Breast Cancer, and Heart Failure Prediction, the framework reached areas under the receiver operating characteristic curve of 0.794, 0.984, and 0.918, respectively. Probabilistic reliability was assessed with a leakage-safe cross-validated calibration analysis reporting Brier score and calibration intercept and slope, and post hoc beta calibration improved probability calibration across datasets. These results indicate that an interpretable, statistically motivated framework can perform comparably to more complex models while providing explicit decision rules expressed in clinical units and risk estimates that are reliable after calibration. A complete worked example further shows that model inference can be reproduced from a printed reference table using only basic arithmetic, without software or proprietary tools, a property that may support auditable use of artificial intelligence in clinical settings.
The authors' abstract, as published at the source. Bioengineering, 2026 · DOI ↗
Continue with a free account
Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.
Continue free on the webSign in with Google or Apple; no card needed. You come back to this paper.
On your phone:
Field: Health Information Management
Health Information ManagementHealth Professions