PofoliaShared via Pofolia

Scientific Reports· 2026Q1

Cluster-aware imputation: a hybrid XGBoost-based model for missing data

Mahmoud M. Ismail, Mohamed Emad, Mai Mohamed

Short summary

A novel hybrid model, ClusteringImputer, uses PCA, K-Means, and cluster-specific XGBoost regressors to impute missing data, achieving O(N) complexity and better accuracy than KNN on large, mixed-type datasets.

AI-generated from the title and abstract; the full text is not read.

Key points

  • Introduces ClusteringImputer, a hybrid model integrating PCA, K-Means, and cluster-specific XGBoost for missing data imputation.
  • The model partitions feature space into subgroups to capture local correlation structures.
  • Achieves linear O(N) computational complexity, significantly outperforming KNN's O(N^2) scaling.
  • Demonstrates superior accuracy and scalability on large (N > 10^4), mixed-type datasets compared to instance-based methods.

AI-generated from the title and abstract; the full text is not read.

Abstract

Abstract The fast growth of high-dimensional datasets in computational biology, industrial IoT, and socio-economic research has made the issue of missing data more challenging. This common problem reduces statistical power and introduces systematic bias in later machine learning tasks. Traditional methods, such as listwise deletion and univariate mean substitution, fail to keep the multivariate covariance structures that are required for accurate decisions. In contrast to traditional Multivariate Imputation by Chained Equations (MICE) frameworks whereby only global regression models that are based on the assumption of single data distribution are mostly used, the modern MICE framework provides improvements through inter-variable modeling. This assumption is not very accurate when it comes to hidden subpopulations and local variances in real datasets. The present work introduces a novel approach in the form of the Cluster Imputation Framework (ClusteringImputer) that is a hybrid model which incorporates unsupervised dimensionality reduction and density-based clustering with supervised gradient boosting. This method specifically captures the local correlation structures by first partitioning the feature space into similar subgroups using Principal Component Analysis and K-Means clustering and then training cluster-specific XGBoost regressors. We accompany a thorough benchmarking analysis of 97 datasets from the UCI Machine Learning Repository; thus, our framework is proven to achieve asymptotic linearity of $$O\left(N\right)$$ in computational complexity. This stands in strong opposition to the K-Nearest Neighbor's (KNN) quadratic $$O\left({N}^{2}\right)$$ scaling. The findings show that despite the fact that instance-based methods work well on continuous low-dimensional manifolds, the Cluster-Aware approach is more efficient with better reconstruction accuracy and scalability in the large scale with more than $$\left(N>{10}^{4}\right)$$ instances, particularly in mixed-type and categorical-dominant situations.

The authors' abstract, as published at the source. Scientific Reports, 2026 · DOI ↗

TakeawaysPremium
Ask the paperFree account

Continue with a free account

Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.

Continue free on the web

Sign in with Google or Apple; no card needed. You come back to this paper.

On your phone:

Field: Statistics and Probability

Statistics and ProbabilityMathematics