PofoliaShared via Pofolia

BMC Medical Informatics and Decision Making· 2026Q1

CleanSurvival: automated data preprocessing for time-to-event models using reinforcement learning

Yousef Koka, David Selby, Gerrit Großmann, Kathan Pandya et al.

Short summary

CleanSurvival, a reinforcement-learning framework, automates data preprocessing for time-to-event (survival) models, improving predictive performance over baselines on real-world datasets.

AI-generated from the title and abstract; the full text is not read.

Key points

  • CleanSurvival uses Q-learning to automate data preprocessing for survival analysis.
  • The framework handles continuous and categorical variables, optimizing imputation, outlier detection, and feature extraction.
  • Experiments show improved predictive performance on real-world datasets compared to simple baselines.
  • A simulation study confirms effectiveness across different types and levels of missingness and noise.

AI-generated from the title and abstract; the full text is not read.

Abstract

Abstract Background Data preprocessing is often paid little attention in machine learning, despite its potentially significant impact on model performance. While automated machine learning pipelines are starting to recognise and integrate data preprocessing into their solutions for classification and regression tasks, this integration is lacking for more specialised tasks like time-to-event models for censored data. As a result, survival analysis not only faces the general challenges of data preprocessing but also suffers from the lack of tailored, automated solutions in this area. Method To address this gap, this paper presents , a reinforcement-learning-based solution for optimizing preprocessing pipelines, extended specifically for survival analysis. The framework can handle continuous and categorical variables. It builds upon Learn2Clean’s $$Q$$ -learning to select which combination of data imputation, outlier detection and feature extraction techniques achieves optimal performance for a Cox, random forest, neural network or user-supplied time-to-event model. The Python package is available on GitHub: https://github.com/datasciapps/CleanSurvival . Results Experimental benchmarks on real-world datasets show that the $$Q$$ -learning-based data preprocessing can improve predictive performance relative to simple baselines, while runtime behaviour is condition-dependent and most clearly interpretable in the best-covered benchmark cells. Furthermore, a simulation study demonstrates effectiveness across different types and levels of missingness and noise. Conclusion With an increase in the use of machine learning, it becomes important to generalise AutoML pipelines to a variety of models now present, including survival analysis. Tools like , which integrate preprocessing for survival analysis, can make survival studies faster and easier to perform, while also yielding more robust results.

The authors' abstract, as published at the source. BMC Medical Informatics and Decision Making, 2026 · DOI ↗

TakeawaysPremium
Ask the paperFree account

Continue with a free account

Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.

Continue free on the web

Sign in with Google or Apple; no card needed. You come back to this paper.

On your phone:

Field: Management Science and Operations Research

Management Science and Operations ResearchDecision Sciences