Diagnostics· 2026Q2
Non-Invasive Prediction of Metabolic Syndrome Using Explainable Machine Learning
- 0citations
- Q2SCImago
- 2026year
Short summary
Explainable machine learning models using non-invasive data achieved moderate discrimination for metabolic syndrome (MetS), with Ensemble and Random Forest models showing the highest AUC of 0.751 and 0.745, respectively.
AI-generated from the title and abstract; the full text is not read.
Key points
- Developed explainable ML models for MetS prediction using non-invasive data (demographics, anthropometrics, behavior).
- Ensemble and Random Forest models showed the highest predictive performance with AUCs of 0.751 and 0.745, respectively.
- Key predictors identified by SHAP analysis include BMI, age, and body weight, along with physical activity and dietary/sociodemographic variables.
- Model performance remained similar in a balanced sensitivity cohort, suggesting robustness.
AI-generated from the title and abstract; the full text is not read.
Abstract
Background: Metabolic syndrome (MetS) is a complex health problem significantly associated with cardiovascular diseases and type 2 diabetes mellitus. Traditional diagnostic approaches rely on invasive biochemical markers, which limit their accessibility. Here, we developed an explainable machine learning (ML) framework for MetS prediction based on non-invasive demographic, anthropometric, and behavioral variables. Method: We conducted a retrospective observational study that included 1090 participants with MetS and 584 without MetS, using six supervised ML models, including Random Forest, AdaBoost, K-Nearest Neighbors, Bagging, Logistic Regression, and Ensemble learning. We trained and validated these models using stratified 10-fold cross-validation. Predictive performance was evaluated using accuracy, precision, recall, specificity, negative predictive value, area under the receiver operating characteristic curve (AUC) and F1-score. Shapley Additive Explanations (SHAP) applied to the Random Forest classifier evaluated interpretability. A 1:1 balanced sensitivity analysis was subsequently performed using the same leakage-safe predictor set and validation framework to assess whether model performance was materially influenced by outcome prevalence. Results: In the natural-prevalence primary cohort (n = 1674; 65.1% with MetS), leakage-safe models showed moderate discrimination. Ensemble achieved the highest mean AUC (0.751, 95% CI 0.731–0.771), followed by Random Forest (0.745, 95% CI 0.723–0.767) and Bagging (0.742, 95% CI 0.725–0.760). SHAP analysis identified (body mass index) BMI, age, and body weight as the dominant contributors, with additional contributions from physical activity, and dietary and sociodemographic variables. In the 1:1 balanced sensitivity cohort (n = 1000), discrimination remained similar, with Ensemble achieving the highest mean AUC (0.742, 95% CI 0.704–0.780). Conclusion: Explainable ML models based on accessible non-laboratory variables provided moderate discrimination of prevalent MetS and may serve as candidate pre-laboratory risk-stratification tools. They should not replace established diagnostic criteria, and external validation, recalibration, and prospective evaluation of clinically relevant thresholds are required before implementation.
The authors' abstract, as published at the source. Diagnostics, 2026 · DOI ↗
The rest is in the Pofolia app
Takeaways and questions to the paper; new summaries every day for your field. Free.
Sign in on the web to openField: Health Information Management
Health Information ManagementHealth Professions