Scientific Reports· 2026Q1
Evaluating the performance of ensemble learning methods in diabetes disease classification
- 0citations
- Q1SCImago
- 2026year
Short summary
Ensemble learning methods (Bagging, Boosting, Stacking) combined with SMOTE for class imbalance achieved high accuracy in diabetes classification, with Light Gradient Boosting and Bagging showing top performance (up to 99.09%) on benchmark datasets.
AI-generated from the title and abstract; the full text is not read.
Key points
- Ensemble learning methods (Bagging, Boosting, Stacking) were evaluated for diabetes classification on three benchmark datasets.
- SMOTE was used to address class imbalance, applied after train-test splitting and within cross-validation folds.
- Performance was assessed using accuracy, precision, recall, F1-score, ROC-AUC, and calibration.
- Top accuracies achieved were 75.97% (PID, Light Gradient Boosting), 98.50% (Frankfurt, Bagging/Light Gradient Boosting), and 99.09% (Sylhet, Bagging).
AI-generated from the title and abstract; the full text is not read.
Abstract
Abstract Diabetes mellitus is a widespread metabolic disorder marked by chronic hyperglycemia and severe complications. Early and accurate detection is crucial for effective management and preventing disease progression. This study systematically evaluates the performance of three ensemble learning strategies Bagging, Boosting, and Stacking on three benchmark diabetes datasets: Pima Indians Diabetes (PID), Frankfurt Hospital Diabetes, and Sylhet Hospital Diabetes. To address class imbalance while preventing data leakage, the Synthetic Minority Oversampling Technique (SMOTE) was applied exclusively to the training data after train-test splitting and independently within each cross-validation fold. Furthermore, all models were evaluated using stratified k-fold cross-validation, and performance was assessed using accuracy, precision, recall, F1-score, receiver operating characteristic area under the curve (ROC-AUC), and calibration analysis. Statistical significance testing was additionally conducted to compare the performance of competing ensemble methods. Experimental results show that all three ensemble paradigms achieved strong performance after SMOTE, with the best-performing model varying by dataset rather than one paradigm uniformly dominating. On the PID dataset, Light Gradient Boosting achieved the highest accuracy (75.97%) On the Frankfurt dataset Bagging and Light Gradient Boosting reached the highest accuracy (98.50%), while on the Sylhet dataset, Bagging perfect accuracy (99.09%) closely followed by Random Forest, Extra Trees and Gradient Boosting (99.03%). These findings underscore the effectiveness of combining SMOTE with Boosting-based ensembles to mitigate class imbalance and improve diabetes classification, highlighting the critical role of both data preprocessing and algorithm selection in achieving high predictive performance for precision medicine and clinical decision support.
The authors' abstract, as published at the source. Scientific Reports, 2026 · DOI ↗
Continue with a free account
Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.
Continue free on the webSign in with Google or Apple; no card needed. You come back to this paper.
On your phone:
Field: Health Information Management
Health Information ManagementHealth Professions