PofoliaShared via Pofolia

Computational Statistics· 2026Q2

Regularisation of regression trees by summation of p values

Nils Engler, Mathias Lindholm, Filip Lindskog, Taariq Nazar

Short summary

A new deterministic, in-sample method for stopping the growth of CART regression trees uses node-wise p-value summation, offering an alternative to time-consuming cross-validation.

AI-generated from the title and abstract; the full text is not read.

Key points

  • Proposes a deterministic, in-sample method for stopping CART regression tree growth using node-wise p-value summation.
  • Method is derived from change point detection, with the null hypothesis being no signal.
  • Allows for arbitrary covariate dimensions and provides an upper bound for the p-value of an entire tree.
  • Demonstrates high probability of detecting signals with sufficient sample size.
  • Illustrates application in constructing deterministic, auto-calibrated predictors.

AI-generated from the title and abstract; the full text is not read.

Abstract

Abstract The standard procedure to decide on the complexity of a CART regression tree is to use cross-validation with the aim of obtaining a predictor that generalises well to unseen data. The randomness in the selection of folds implies that the selected CART regression tree is not a deterministic function of the data. Moreover, the cross-validation procedure may become time consuming and result in inefficient use of training data. We propose a simple deterministic in-sample method that can be used for stopping the growing of a CART regression tree based on node-wise statistical tests. This testing procedure is derived using a connection to change point detection, where the null hypothesis corresponds to no signal. The suggested p value based procedure allows us to consider covariate vectors of arbitrary dimension and allows us to bound the p value of an entire tree from above. Further, we show that the test detects a not too weak signal with a high probability, given a not too small sample size. We illustrate our methodology and the asymptotic results on both simulated and real world data. Additionally, we illustrate how the p value based method can be used to construct a deterministic piece-wise constant auto-calibrated predictor based on a given black-box predictor.

The authors' abstract, as published at the source. Computational Statistics, 2026 · DOI ↗

TakeawaysIn the app
Ask the paperIn the app

The rest is in the Pofolia app

Takeaways and questions to the paper; new summaries every day for your field. Free.

Sign in on the web to open

Field: Computational Theory and Mathematics

Computational Theory and MathematicsComputer Science