PofoliaShared via Pofolia

SIAM Journal on Mathematics of Data Science· 2026Q1

High-Dimensional Analysis of Ridge Regression for Non-identically Distributed Data with a Variance Profile

Jérémie Bigot, Issa-Mbenard Dabo, Camille Male

Short summary

This paper analyzes the ridge regression estimator for high-dimensional, independent but non-identically distributed data, deriving deterministic equivalents for predictive risk and degrees of freedom.

AI-generated from the title and abstract; the full text is not read.

Key points

  • Analyzes ridge regression for high-dimensional, independent non-identically distributed data with a random matrix predictor and variance profile.
  • Derives deterministic equivalents for the predictive risk and degrees of freedom of the ridge estimator.
  • Highlights the emergence of double descent for minimum-norm least-squares estimators under specific variance profiles.
  • Identifies variance profiles for which predictive risk deviates from the double descent pattern.

AI-generated from the title and abstract; the full text is not read.

Abstract

Abstract. High-dimensional linear regression has been thoroughly studied in the context of independent and identically distributed data. We propose to investigate high-dimensional regression models for independent but non-identically distributed data. To this end, we suppose that the set of observed predictors (or features) is a random matrix with a variance profile and with dimensions growing at a proportional rate. Assuming a random effect model, we study the predictive risk of the ridge estimator for linear regression with such a variance profile. In this setting, we provide deterministic equivalents of this risk and of the degree of freedom of the ridge estimator. For a certain class of variance profile, our work highlights the emergence of the well-known double descent phenomenon in high-dimensional regression for the minimum-norm least-squares estimator when the ridge regularization parameter goes to zero. We also exhibit variance profiles for which the shape of this predictive risk differs from double descent. The proofs of our results are based on tools from random matrix theory in the presence of a variance profile that have not been considered so far to study regression models. Numerical experiments are provided to show the accuracy of the aforementioned deterministic equivalents on the computation of the predictive risk of ridge regression. We also investigate the similarities and differences that exist with the standard setting of independent and identically distributed data.

The authors' abstract, as published at the source. SIAM Journal on Mathematics of Data Science, 2026 · DOI ↗

TakeawaysPremium
Ask the paperFree account

Continue with a free account

Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.

Continue free on the web

Sign in with Google or Apple; no card needed. You come back to this paper.

On your phone:

Field: Computational Theory and Mathematics

Computational Theory and MathematicsComputer Science