Scientific Reports· 2026Q1
Küme farkındalığına sahip atama: eksik veriler için hibrit bir XGBoost tabanlı model
Cluster-aware imputation: a hybrid XGBoost-based model for missing data
- 0atıf
- Q1SCImago
- 2026yıl
Kısa özet
PCA, K-Means ve küme özgü XGBoost regressor'larını kullanan yeni bir hibrit model olan ClusteringImputer, eksik verileri atamak için O(N) karmaşıklık ve büyük, karma türdeki veri kümelerinde KNN'den daha iyi doğruluk elde eder.
Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.
Ana noktalar
- Eksik veri ataması için PCA, K-Means ve küme özgü XGBoost'u entegre eden hibrit bir model olan ClusteringImputer'ı tanıtır.
- Model, yerel korelasyon yapısını yakalamak için özellik uzayını alt gruplara ayırır.
- KNN'nin O(N^2) ölçeklenmesini önemli ölçüde aşan doğrusal O(N) hesaplama karmaşıklığına ulaşır.
- Büyük (N > 10^4), karma türdeki veri kümelerinde örnek tabanlı yöntemlere kıyasla üstün doğruluk ve ölçeklenebilirlik gösterir.
Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.
Özet (abstract)
Abstract The fast growth of high-dimensional datasets in computational biology, industrial IoT, and socio-economic research has made the issue of missing data more challenging. This common problem reduces statistical power and introduces systematic bias in later machine learning tasks. Traditional methods, such as listwise deletion and univariate mean substitution, fail to keep the multivariate covariance structures that are required for accurate decisions. In contrast to traditional Multivariate Imputation by Chained Equations (MICE) frameworks whereby only global regression models that are based on the assumption of single data distribution are mostly used, the modern MICE framework provides improvements through inter-variable modeling. This assumption is not very accurate when it comes to hidden subpopulations and local variances in real datasets. The present work introduces a novel approach in the form of the Cluster Imputation Framework (ClusteringImputer) that is a hybrid model which incorporates unsupervised dimensionality reduction and density-based clustering with supervised gradient boosting. This method specifically captures the local correlation structures by first partitioning the feature space into similar subgroups using Principal Component Analysis and K-Means clustering and then training cluster-specific XGBoost regressors. We accompany a thorough benchmarking analysis of 97 datasets from the UCI Machine Learning Repository; thus, our framework is proven to achieve asymptotic linearity of $$O\left(N\right)$$ in computational complexity. This stands in strong opposition to the K-Nearest Neighbor's (KNN) quadratic $$O\left({N}^{2}\right)$$ scaling. The findings show that despite the fact that instance-based methods work well on continuous low-dimensional manifolds, the Cluster-Aware approach is more efficient with better reconstruction accuracy and scalability in the large scale with more than $$\left(N>{10}^{4}\right)$$ instances, particularly in mixed-type and categorical-dominant situations.
Yazarların özeti; kaynağından alınmıştır. Scientific Reports, 2026 · DOI ↗
Ücretsiz hesapla devam et
Makaleye Sor ile bu makaleye günde 3 soru ücretsiz; makaleyi kaydet, kaynakçasını al, ilgi alanına göre her gün yeni özetler. Çıkarımlar Premium.
Web'de ücretsiz devam etGoogle ya da Apple hesabınla giriş; kart istemez. Bu makaleye geri dönersin.
Telefonda:
Alan: İstatistik ve Olasılık
Statistics and ProbabilityMathematics