Journal Of Big Data· 2025Q1
Accuracy, precision, recall, f1-score, or MCC? empirical evidence from advanced statistics, ML, and XAI for evaluating business predictive models
- 95citations
- Q1SCImago
- 2025year
Short summary
The F1-score is the most stable and balanced metric for evaluating imbalanced business classification tasks, outperforming accuracy and precision, with MCC offering complementary value.
AI-generated from the title and abstract; the full text is not read.
Key points
- F1-score provides the most stable and balanced evaluation for imbalanced business classification tasks.
- MCC offers complementary diagnostic value to the F1-score.
- Accuracy and precision demonstrate limited robustness under class imbalance.
- A novel 3D metric-conditioned SHAP analysis links feature contributions to classification thresholds and evaluation metrics.
AI-generated from the title and abstract; the full text is not read.
Abstract
Imbalanced datasets pose a persistent challenge in business data mining, particularly in high-stakes domains such as financial risk prediction and customer churn analysis, where the minority class often carries disproportionate operational and financial consequences. Although widely used evaluation metrics–such as accuracy, precision, recall, F1-score, and Matthews Correlation Coefficient (MCC)–are commonly applied in practice, there remains no empirical consensus on which metric offers the most reliable performance under real-world conditions. Existing studies lack a unified, statistically validated framework that accounts for threshold sensitivity, input noise, and interpretability–factors critical to business decision-making. To address this gap, we present a comprehensive and statistically rigorous evaluation of performance metrics for imbalanced business classification tasks. Using two benchmark datasets with distinct sizes and imbalance ratios–the Default of Credit Card Clients dataset and the Telco Customer Churn dataset–we evaluate five commonly used machine learning models: Logistic Regression (LR), Decision Tree (DT), Random Forest (RF), Extreme Gradient Boosting (XGBoost), and k-Nearest Neighbors (KNN). Our methodology incorporates static and dynamic threshold analysis, Gaussian noise robustness testing, bootstrap confidence intervals, McNemar’s test, Cohen’s kappa, and analysis of variance (ANOVA) to assess the statistical reliability of performance metrics. In addition, we introduce a novel two-stage explainable artificial intelligence (XAI) framework using SHapley Additive exPlanations (SHAP). The first stage employs standard SHAP visualizations (bar and beeswarm plots) to ensure baseline interpretability. The second stage extends this with a novel 3D metric-conditioned SHAP analysis, linking feature contributions to variations in classification thresholds and evaluation metrics. Our findings show that the F1-score consistently provides the most stable and balanced evaluation across datasets and testing conditions, with MCC offering complementary diagnostic value. In contrast, accuracy and precision demonstrate limited robustness under class imbalance. By combining statistical rigor with interpretable AI, this study offers the most comprehensive guidance to date for selecting performance metrics in imbalanced business classification, with practical implications for model deployment in finance, marketing, and customer analytics.
The authors' abstract, as published at the source. Journal Of Big Data, 2025 · DOI ↗
Continue with a free account
Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.
Continue free on the webSign in with Google or Apple; no card needed. You come back to this paper.
On your phone:
Field: Artificial Intelligence
Artificial IntelligenceComputer Science