Scientific Reports· 2026Q1
Diyabet hastalığı sınıflandırmasında topluluk öğrenme yöntemlerinin performansının değerlendirilmesi
Evaluating the performance of ensemble learning methods in diabetes disease classification
- 0atıf
- Q1SCImago
- 2026yıl
Kısa özet
Sınıf dengesizliği için SMOTE ile birleştirilen topluluk öğrenme yöntemleri (Bagging, Boosting, Stacking), yüksek doğruluk oranları elde etti; Light Gradient Boosting ve Bagging, referans veri setlerinde en iyi performansı ( %99.09'a kadar) gösterdi.
Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.
Ana noktalar
- Diyabet sınıflandırması için üç referans veri setinde topluluk öğrenme yöntemleri (Bagging, Boosting, Stacking) değerlendirildi.
- Sınıf dengesizliğini ele almak için, eğitim-test ayırmasından sonra ve çapraz doğrulama katları içinde SMOTE uygulandı.
- Performans doğruluk, kesinlik, duyarlılık, F1-skoru, ROC-AUC ve kalibrasyon analizleri kullanılarak değerlendirildi.
- Elde edilen en yüksek doğruluk oranları şunlardır: %75.97 (PID, Light Gradient Boosting), %98.50 (Frankfurt, Bagging/Light Gradient Boosting) ve %99.09 (Sylhet, Bagging).
Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.
Özet (abstract)
Abstract Diabetes mellitus is a widespread metabolic disorder marked by chronic hyperglycemia and severe complications. Early and accurate detection is crucial for effective management and preventing disease progression. This study systematically evaluates the performance of three ensemble learning strategies Bagging, Boosting, and Stacking on three benchmark diabetes datasets: Pima Indians Diabetes (PID), Frankfurt Hospital Diabetes, and Sylhet Hospital Diabetes. To address class imbalance while preventing data leakage, the Synthetic Minority Oversampling Technique (SMOTE) was applied exclusively to the training data after train-test splitting and independently within each cross-validation fold. Furthermore, all models were evaluated using stratified k-fold cross-validation, and performance was assessed using accuracy, precision, recall, F1-score, receiver operating characteristic area under the curve (ROC-AUC), and calibration analysis. Statistical significance testing was additionally conducted to compare the performance of competing ensemble methods. Experimental results show that all three ensemble paradigms achieved strong performance after SMOTE, with the best-performing model varying by dataset rather than one paradigm uniformly dominating. On the PID dataset, Light Gradient Boosting achieved the highest accuracy (75.97%) On the Frankfurt dataset Bagging and Light Gradient Boosting reached the highest accuracy (98.50%), while on the Sylhet dataset, Bagging perfect accuracy (99.09%) closely followed by Random Forest, Extra Trees and Gradient Boosting (99.03%). These findings underscore the effectiveness of combining SMOTE with Boosting-based ensembles to mitigate class imbalance and improve diabetes classification, highlighting the critical role of both data preprocessing and algorithm selection in achieving high predictive performance for precision medicine and clinical decision support.
Yazarların özeti; kaynağından alınmıştır. Scientific Reports, 2026 · DOI ↗
Ücretsiz hesapla devam et
Makaleye Sor ile bu makaleye günde 3 soru ücretsiz; makaleyi kaydet, kaynakçasını al, ilgi alanına göre her gün yeni özetler. Çıkarımlar Premium.
Web'de ücretsiz devam etGoogle ya da Apple hesabınla giriş; kart istemez. Bu makaleye geri dönersin.
Telefonda:
Alan: Sağlık Bilgi Yönetimi
Health Information ManagementHealth Professions