Diagnostics· 2026Q2
Explainable Machine Learning for Prediction of Future High-Severity States Using Longitudinal APACHE-II Trajectories: Internal Validation and Cross-Cohort Portability Assessment
- 0citations
- Q2SCImago
- 2026year
Short summary
A CatBoost machine learning model using longitudinal APACHE-II scores predicted subsequent high-severity states (APACHE-II ≥ 20) in ICU patients with 0.921 ROC-AUC on independent testing, significantly outperforming single-score or simpler trajectory models.
AI-generated from the title and abstract; the full text is not read.
Key points
- A CatBoost model using longitudinal APACHE-II data achieved 0.921 ROC-AUC for predicting high-severity states (APACHE-II ≥ 20) in ICU patients.
- The longitudinal model significantly outperformed single APACHE-II scores (ROC-AUC 0.553) and simpler trajectory calculations (ROC-AUC 0.633).
- SHAP analysis indicated that multiple APACHE-II observations and derived trajectory descriptors were important predictors.
- Adding laboratory data or baseline characteristics did not improve prediction accuracy over longitudinal APACHE-II data alone.
AI-generated from the title and abstract; the full text is not read.
Abstract
Background: Longitudinal severity assessment may provide more informative risk stratification than reliance on a single admission score in intensive care. This study developed and internally validated an explainable machine learning framework for predicting a subsequent APACHE-II-defined high-severity state using repeated APACHE-II assessments in intensive care unit (ICU) patients receiving total parenteral nutrition (TPN). The endpoint was defined as an APACHE-II score of ≥ 20 at the final assessment (T10), while predictor information was restricted to measurements obtained at T0–T9. Methods: A retrospective institutional cohort of 844 adult ICU patients was analyzed, of whom 214 (25.4%) met the predefined high-severity endpoint. Four principal feature representations were evaluated: admission APACHE-II, longitudinal APACHE-II information combining the raw T0–T9 sequence with derived trajectory descriptors, longitudinal laboratory information, and multimodal combinations including baseline characteristics. Six machine learning algorithms were compared using stratified five-fold cross-validation and independent hold-out testing. Additional analyses included simpler APACHE-II comparators, sequential observation truncation, nested cross-validation, calibration assessment, threshold sensitivity analysis, and SHAP-based model interpretation. Because equivalent longitudinal APACHE-II measurements and an equivalent endpoint were unavailable in eICU, a separate harmonized XGBoost model was used only for exploratory cross-cohort portability assessment. Results: The longitudinal APACHE-II CatBoost model achieved the highest performance, with a cross-validated ROC-AUC of 0.924 and an independent hold-out ROC-AUC of 0.921 (95% CI: 0.878–0.960). Hold-out calibration was good (Brier score = 0.095, calibration intercept = − 0.03, slope = 0.98), and nested cross-validation yielded a pooled out-of-fold ROC-AUC of 0.908 and a mean outer-fold ROC-AUC of 0.920 ± 0.038. The latest APACHE-II assessment alone (T9; ROC-AUC = 0.553), change from baseline (0.576), ordinal slope (0.539), and logistic regression using engineered APACHE-II descriptors (0.633) performed substantially below the full longitudinal CatBoost model. Sequential truncation showed that discrimination was retained after removing later observations, although performance varied non-monotonically across truncated sequences. SHAP analysis showed that multiple APACHE-II observations contributed prominently to model predictions, with trajectory-derived descriptors providing complementary contributions. Adding laboratory trajectories or baseline characteristics did not improve discrimination over the longitudinal APACHE-II representation. Conclusions: Nonlinear modeling of repeated APACHE-II assessments provided strong internal discrimination of a subsequent APACHE-II-defined high-severity state in this TPN-treated ICU cohort and substantially outperformed single-score and simpler longitudinal comparators. However, the predictors and endpoint share the APACHE-II construct, observation indices were sequential rather than standardized clock-time intervals, and the primary model could not undergo conventional external validation in eICU. These findings therefore support the internal predictive value of longitudinal APACHE-II modeling for this specific severity-state task and warrant prospective multicenter evaluation using standardized timing and independent clinical outcomes.
The authors' abstract, as published at the source. Diagnostics, 2026 · DOI ↗
Continue with a free account
Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.
Continue free on the webSign in with Google or Apple; no card needed. You come back to this paper.
On your phone:
Field: Nutrition and Dietetics
Nutrition and DieteticsNursing