Robots Atlas>ROBOTS ATLAS
Training

Ridge

1970ActivePublished: 24 August 2026Updated: 24 August 2026Published
Key innovation
Adding an L2 penalty on coefficient magnitude to linear regression, yielding stable, biased estimators even under multicollinearity and a near-singular XᵀX matrix.
Category
Training
Abstraction level
Building block
Operation level
TrainingModel
Use cases
Regression with collinear featuresIll-posed and inverse problemsHigh-dimensional data (p > n)Stabilizing linear modelsSmoothing in statistics and econometricsMachine-learning baseline

How it works

An α·Σwⱼ² penalty term is added to the least-squares loss. Optimizing this objective yields modified normal equations whose closed-form solution is ŵ = (XᵀX + αI)⁻¹Xᵀy. Adding αI raises a 'ridge' (hence the name) along the diagonal of XᵀX, guaranteeing invertibility and improved conditioning. The parameter α governs the tradeoff: as α→0 the estimator approaches OLS, while large α shrinks coefficients strongly toward zero. Features should be standardized because the penalty is scale-dependent; the intercept is usually left unpenalized. The optimal α is selected via cross-validation (e.g. RidgeCV, generalized cross-validation / GCV).

Problem solved

Ordinary least squares (OLS) produces high-variance estimates when features are strongly correlated (multicollinearity) or when XᵀX is near-singular or non-invertible (e.g. when features outnumber observations), leading to unstable, overfitted models. Ridge addresses this via L2 regularization, which stabilizes the matrix inversion and curbs overfitting.

Components

Linear modelPredictor

A parametric linear predictor ŷ = Xw whose coefficients are estimated under regularization.

INFeature matrix: n observations, d features.
OUTPrediction vector.
L2 penalty termRegularizer

A regularization term penalizing the squared Euclidean norm of the coefficients, shrinking them toward zero.

L1 penalty (Lasso)L1 penalty yielding sparse solutions (coefficient zeroing).
Elastic NetA combination of L1 and L2 penalties.

Official

Closed-form estimatorEstimator

Solution of the modified normal equations ŵ = (XᵀX + αI)⁻¹Xᵀy.

Official

Implementation

Implementation pitfalls
Skipping feature standardizationHigh

The L2 penalty is scale-dependent; without standardization, large-scale features are under-regularized.

Fix:Standardize features (zero mean, unit variance) before fitting.
Penalizing the interceptMedium

Including the intercept in the penalty introduces a target-shift-dependent error.

Fix:Exclude the intercept from regularization.
Poor choice of αHigh

Too-small α fails to curb overfitting; too-large α causes underfitting.

Fix:Select α via cross-validation (RidgeCV / GCV).

Evolution

Original paper · 1970 · Technometrics, 12(1), 55–67 · Arthur E. Hoerl
Ridge Regression: Biased Estimation for Nonorthogonal Problems
Arthur E. Hoerl, Robert W. Kennard
1963
Tikhonov regularization
Inflection point

Andrey Tikhonov formulates the regularization method for ill-posed problems — the mathematical equivalent of Ridge.

1970
Hoerl & Kennard formalize Ridge Regression
Inflection point

Two Technometrics papers introduce ridge regression in statistics and its applications to nonorthogonal problems.

1996
Lasso (L1 penalty) as an alternative

Robert Tibshirani introduces the Lasso, yielding sparse solutions — a contrast to Ridge's proportional shrinkage.

2005
Elastic Net combines L1 and L2

Zou and Hastie combine L1 and L2 penalties, reconciling Lasso sparsity with Ridge stability.

Hyperparameters (configurable axes)

Regularization strength (α / λ)Critical

Non-negative L2 penalty coefficient; controls the bias–variance tradeoff. α→0 approaches OLS, large α strongly shrinks coefficients.

0.1 – 1.0Typical starting range with standardized features.
logspace(-6, 6)Grid for cross-validation (RidgeCV).
SolverMedium

Solution algorithm: SVD/Cholesky decomposition, lsqr, sparse_cg, sag/saga, lbfgs.

Feature standardizationHigh

Ridge is not scale-invariant; standardization is usually required for meaningful regularization.

Fit interceptLow

Whether to fit and (usually) leave unpenalized the intercept.

Computational complexity

Time complexity: O(n · d²). Space complexity: O(d²).

Compute bottleneck

Inversion / factorization of (XᵀX + αI)

The main cost is solving the normal-equations system with the regularized d×d matrix.

Execution paradigm

Primary mode
Dense

All coefficients are active in every prediction; no routing or conditional activation.

Activation pattern
All paths active

Parallelism

Parallelism level
Fully parallel

Dense linear-algebra operations (matrix multiply, factorization, dot product at prediction time) parallelize well in BLAS/LAPACK libraries.

Scope
TrainingInference

Hardware requirements

Primary

Built on basic linear algebra, it runs on any hardware.

Good fit

For small-to-medium datasets, a CPU with BLAS/AVX is fully sufficient.

Possible

A GPU accelerates matrix operations on very large data.