PofoliaPofolia ile paylaşıldı

Journal Of Big Data· 2025Q1

Doğruluk, kesinlik, duyarlılık, f1-skoru mu, yoksa MCC mi? İşletme tahmin modellerini değerlendirmek için ileri istatistik, ML ve XAI'den ampirik kanıtlar

Accuracy, precision, recall, f1-score, or MCC? empirical evidence from advanced statistics, ML, and XAI for evaluating business predictive models

Khaled Mahmud Sujon, rohayanti binti hassan, Kwonhue Choi, Md Abdus Samad

Kısa özet

Dengesiz iş sınıflandırma görevleri için en kararlı ve dengeli metrik F1-skorudur; doğruluk ve kesinlikten daha iyi performans gösterir ve MCC tamamlayıcı değer sunar.

Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.

Ana noktalar

  • F1-skoru, dengesiz iş sınıflandırma görevleri için en kararlı ve dengeli değerlendirmeyi sağlar.
  • MCC, F1-skoruna tamamlayıcı tanısal değer sunar.
  • Doğruluk ve kesinlik, sınıf dengesizliği altında sınırlı sağlamlık gösterir.
  • Yeni bir 3B metrik koşullu SHAP analizi, özellik katkılarını sınıflandırma eşiklerine ve performans metriklerine bağlar.

Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.

Özet (abstract)

Imbalanced datasets pose a persistent challenge in business data mining, particularly in high-stakes domains such as financial risk prediction and customer churn analysis, where the minority class often carries disproportionate operational and financial consequences. Although widely used evaluation metrics–such as accuracy, precision, recall, F1-score, and Matthews Correlation Coefficient (MCC)–are commonly applied in practice, there remains no empirical consensus on which metric offers the most reliable performance under real-world conditions. Existing studies lack a unified, statistically validated framework that accounts for threshold sensitivity, input noise, and interpretability–factors critical to business decision-making. To address this gap, we present a comprehensive and statistically rigorous evaluation of performance metrics for imbalanced business classification tasks. Using two benchmark datasets with distinct sizes and imbalance ratios–the Default of Credit Card Clients dataset and the Telco Customer Churn dataset–we evaluate five commonly used machine learning models: Logistic Regression (LR), Decision Tree (DT), Random Forest (RF), Extreme Gradient Boosting (XGBoost), and k-Nearest Neighbors (KNN). Our methodology incorporates static and dynamic threshold analysis, Gaussian noise robustness testing, bootstrap confidence intervals, McNemar’s test, Cohen’s kappa, and analysis of variance (ANOVA) to assess the statistical reliability of performance metrics. In addition, we introduce a novel two-stage explainable artificial intelligence (XAI) framework using SHapley Additive exPlanations (SHAP). The first stage employs standard SHAP visualizations (bar and beeswarm plots) to ensure baseline interpretability. The second stage extends this with a novel 3D metric-conditioned SHAP analysis, linking feature contributions to variations in classification thresholds and evaluation metrics. Our findings show that the F1-score consistently provides the most stable and balanced evaluation across datasets and testing conditions, with MCC offering complementary diagnostic value. In contrast, accuracy and precision demonstrate limited robustness under class imbalance. By combining statistical rigor with interpretable AI, this study offers the most comprehensive guidance to date for selecting performance metrics in imbalanced business classification, with practical implications for model deployment in finance, marketing, and customer analytics.

Yazarların özeti; kaynağından alınmıştır. Journal Of Big Data, 2025 · DOI ↗

ÇıkarımlarPremium
Makaleye SorÜcretsiz hesapla

Ücretsiz hesapla devam et

Makaleye Sor ile bu makaleye günde 3 soru ücretsiz; makaleyi kaydet, kaynakçasını al, ilgi alanına göre her gün yeni özetler. Çıkarımlar Premium.

Web'de ücretsiz devam et

Google ya da Apple hesabınla giriş; kart istemez. Bu makaleye geri dönersin.

Telefonda:

Alan: Yapay Zeka

Artificial IntelligenceComputer Science