← Methods repository
CONCEPTADVANCEDlast reviewed 2026-08-24 · updated 2026-08-25 · 2 citations

Regularized Regression: LASSO, Ridge, and Elastic Net

A family of penalized linear and generalized linear models that add a penalty term to the ordinary least-squares or maximum-likelihood objective, shrinking coefficient estimates toward zero (ridge/L2) or exactly to zero (LASSO/L1), with elastic net combining both penalties; used in RWE for variable selection and prediction from thousands of claims codes, high-dimensional propensity score construction, risk prediction when predictors outnumber observations, and as a nuisance model inside double-ML and TMLE.

Machine Learning and Predictivemachine-learningpredictionregularizationvariable-selectionhigh-dimensionallassoridgeelastic-net
On this page
Methods reference only. Use primary source citations and local policy before applying this in a study protocol, regulatory submission, payer dossier, or clinical decision.

What regularized regression is and why it matters in RWE

Ordinary least squares (OLS) and maximum-likelihood estimation find the coefficients that minimize the residual objective with no constraint on coefficient magnitude. In high-dimensional real-world evidence settings — where a claims analyst confronts thousands of ICD-10 diagnosis codes, CPT procedure codes, and NDC drug classes as candidate predictors — unconstrained estimators break down in two ways. First, when the number of predictors p approaches or exceeds the number of observations n, OLS has no unique solution and fits noise perfectly, producing wildly inflated coefficient variance. Second, even when n > p, OLS coefficients have high variance under multicollinearity — and ICD code families (diabetes with renal complications, diabetes without) are almost by definition collinear. Regularized regression adds a penalty term to the objective that trades a controlled amount of bias for a substantial reduction in variance, yielding estimates that generalize better to held-out data and remain estimable when p >> n.

The three penalties: geometry and intuition

All three methods minimize the same penalized objective: minimize { sum of squared (or deviance) residuals + lambda * Penalty(beta) }

The penalty type determines the geometry of the solution:

  • Ridge (L2 penalty): Penalty = sum of beta_j squared. The L2 feasible region is a circle in two dimensions (sphere in higher). The optimization finds where the OLS loss ellipsoid first touches this sphere, and because the sphere has no corners, the solution almost never lands exactly at zero. Every coefficient is shrunk proportionally toward zero but retained. Ridge excels when all predictors carry some signal and when correlated code groups (all diabetes-related diagnoses, all hypertension-related procedure codes) should be kept together rather than having one arbitrarily selected and the rest discarded. Ridge also uniquely defines the coefficient vector even when the OLS design matrix is singular — critical for p > n claims analyses.
  • LASSO (L1 penalty, Least Absolute Shrinkage and Selection Operator): Penalty = sum of |beta_j|. The L1 feasible region is a diamond in two dimensions (cross-polytope in higher dimensions) with sharp corners at the coordinate axes. The OLS loss ellipsoid most often first contacts the L1 region at one of these corners, where all but one or a few coordinates are zero. LASSO simultaneously estimates and selects: it produces a sparse model with many exact zeros. In a claims cost model with 4,000 candidate codes, LASSO at a cross-validated lambda typically retains 15 to 30 codes with non-zero coefficients — a compact, communicable predictive fingerprint. The sparsity comes at a cost: with strongly correlated predictors (two nearly identical complication codes), LASSO arbitrarily selects one and zeros the other, making the selected set unstable across bootstrap replicates.
  • Elastic net (L1 plus L2 mixture): Penalty = alpha * sum(|beta_j|) + (1-alpha) * sum(beta_j squared). The elastic net mixes LASSO sparsity with ridge shrinkage of the selected group. When correlated code families compete for inclusion — four closely related diabetes complication codes where the science implies all four should be present — elastic net tends to keep the group together while zeroing truly irrelevant codes. The mixing parameter alpha controls the L1 fraction: alpha = 1 is pure LASSO, alpha = 0 is pure ridge. A value of alpha between 0.5 and 0.8 is a common default for claims code analyses with hierarchically organized ICD families. A further extension, the group lasso, enforces sparsity at the group level — all members of a pre-specified ICD-10 chapter or ATC drug class either enter or leave together — which is clinically natural for hierarchically organized codes.

Standardization is mandatory before penalizing

The penalty term treats all coefficients on the same footing. A coefficient of 1.0 on age (measured in decades, range 3 to 8) receives the same penalty as a coefficient of 1.0 on a binary ever/never flag (range 0 to 1). Without standardization, predictors measured on large scales are penalized less and predictors on small scales are penalized more — an artifact of units, not biology or confounding structure. Both glmnet (R) and sklearn (Python) standardize predictors internally by default and return coefficients on the original scale; always verify this behavior is active and has not been silently disabled in the implementation.

Lambda selection by cross-validation

Lambda (the penalty strength) is the critical tuning parameter. Too small: the penalty is negligible and the estimator reverts to OLS. Too large: every coefficient is shrunk to zero. In practice, lambda is chosen by k-fold cross-validation across a fine grid of values: fit on k-1 folds, predict the held-out fold, compute the loss (mean squared error for linear outcomes, binomial deviance for logistic). Two conventional choices from the resulting cross-validation curve:

  • lambda.min: the lambda minimizing mean CV error — lowest bias, maximum predictive accuracy.
  • lambda.1se: the largest lambda within one standard error of the minimum — sparser model, slightly higher CV error, more parsimonious. Preferred in RWE contexts where parsimony and communication matter more than the last fraction of AUC or R-squared.

The choice between lambda.min and lambda.1se should be pre-specified in the SAP or protocol and justified scientifically, not chosen post-hoc to maximize a preferred model size.

RWE high-dimensional reality: claims code spaces and p >> n settings

The primary motivation for regularized regression in RWE is the claims code landscape. A standard 365-day lookback window in a Medicare FFS or commercial database generates roughly 3,000 to 6,000 unique ICD-10 codes, 1,500 to 3,000 CPT codes, and 800 to 1,500 NDC drug classes per cohort. OLS is not identified at these dimensions; even moderate-sized cohorts face p >> n for subgroup analyses. The coordinate descent algorithm underlying glmnet cycles through predictors without requiring matrix inversion, handles both n > p and p >> n equally, and scales to millions of observations with tens of thousands of predictors on standard hardware.

In the high-dimensional propensity score (hdPS) pipeline, the empirical Bayes selection step plays a role conceptually similar to LASSO: data-adaptively identifying which of thousands of codes most predict treatment assignment. Full LASSO logistic regression on all candidate codes is increasingly used as a direct alternative or complement to hdPS for propensity score construction, with lambda chosen by cross-validation.

For biomarker panels (genomics, proteomics, pharmacogenomics) and linked registry-claims datasets, p >> n is routine. Ridge excels when all biomarkers are expected to contribute (polygenic scores, pathway-level analyses); LASSO when a small number of truly predictive biomarkers is plausible; elastic net when correlated gene families or pathway blocks should be grouped together.

Interpreting the output

A LASSO claims-cost model fit at a cross-validation-chosen lambda retains 18 of 4,000 candidate codes. The coefficient on "prior insulin use" is 0.35 on the log-cost scale.

Formal interpretation

Penalized coefficients are deliberately biased toward zero — this is the mechanism, not a failure mode. The value 0.35 is a shrunken predictive weight, not an unbiased adjusted effect estimate. It is not entitled to a naive confidence interval: the standard errors from a refitted unpenalized OLS on only the selected 18 codes are invalid. They ignore the selection step (which consumed degrees of freedom not reflected in the refitted model) and consistently understate uncertainty — naive SEs after LASSO are known to be anticonservative. Reporting them as if they were unpenalized OLS standard errors is incorrect. Post-selection inference is an active research area (selective inference, PoSI) requiring specialized software; it is not solved by simply refitting OLS on the LASSO-selected set and using the resulting standard errors.

Selection is also unstable: bootstrap the analysis cohort 200 times and the set of 18 codes changes with every replicate. Prior insulin use may appear in 70 to 80 percent of replicates as a stable predictor; the marginal code at position 18 may appear in only 20 to 30 percent. The specific coefficient 0.35 shifts in each bootstrap replicate, reflecting both sampling variation and selection variation.

Practical interpretation

Statistically correct plain English: the model found prior insulin use one of the strongest cost predictors among the 18 codes retained at the selected lambda. Treat the 18 selected codes as a predictive fingerprint — a compact feature set for scoring new patients' cost risk — not as a list of cost drivers in any causal sense. The selection of 18 codes does not imply these are the 18 most important biological or clinical determinants of cost. If the scientific question is whether insulin use causes higher costs compared with an alternative therapy, this LASSO fit does not answer it and should not be reported as if it does.

The inference warning: regularization optimizes prediction, not causal inference

Regularized regression is designed to minimize prediction error. Using LASSO to select confounders and then refitting unpenalized regression on only the selected set is a two-step procedure with invalid inference. Naive SEs from the refitted model are anticonservative and the point estimate for the treatment effect is inconsistent — the selection step induces a form of bias at the variable level that the refitted standard errors cannot recover.

The principled fix is post-double-selection (Belloni, Chernozhukov, and Hansen): run LASSO of the outcome on all candidate confounders; run LASSO of the exposure on all candidate confounders; take the union of the two selected sets; refit unpenalized regression on this union plus the exposure. This produces root-n-consistent, asymptotically normal estimates of the treatment effect even when thousands of candidate confounders are screened. An equivalent and more general approach is double/debiased ML (see parent entry predictive-and-causal-ml-models-rwe), which incorporates this logic in a cross-fit framework with formal efficiency guarantees.

A second critical rule: never penalize the exposure coefficient itself when the goal is causal effect estimation. The exposure must enter the model without penalty; only the potential confounders receive shrinkage. Penalizing the exposure biases the causal estimate toward the null — a form of dilution bias invisible to standard fit metrics.

Pros, cons, and trade-offs

Pros: Estimable when p approaches or exceeds n — the only linear framework that works in truly high-dimensional claims data without separate dimension reduction. Automatic variable selection (LASSO, elastic net) reduces the model to a compact set of predictors — a communication advantage in HEOR submissions. Ridge eliminates instability under multicollinearity; correlated code families are retained collectively rather than having arbitrary members selected. The coordinate descent algorithm (glmnet) is fast enough for millions of observations and tens of thousands of predictors. Extends to other outcomes via penalized GLMs: logistic (readmission, treatment failure), Poisson (utilization counts), Cox proportional hazards (time-to-event), multinomial. Lambda chosen by cross-validation is principled and reproducible.

Cons: Coefficients are biased by construction; they cannot be reported as unbiased adjusted effects without post-double-selection or a causal ML wrapper. Naive confidence intervals after LASSO selection are invalid and anticonservative. LASSO selection is unstable with strongly correlated predictors; elastic net or group lasso is needed when correlated groups must be represented collectively. The lambda.min vs lambda.1se choice materially affects the selected set and should be pre-specified. SAS does not provide a full cross-validated regularization path equivalent to glmnet; PROC GLMSELECT LASSO is available for moderate dimensions but requires additional calibration steps.

When to use

  • High-dimensional prediction with tens to thousands of candidate predictors from a claims lookback window, biomarker panel, or genomic dataset; OLS is not identified or has high variance; the goal is a risk score or cost prediction model for budget impact or patient stratification.
  • Propensity score construction from thousands of code candidates as an alternative to stepwise selection or the hdPS empirical Bayes step, with lambda chosen by CV.
  • Confounder screening: reduce 4,000 codes to a manageable set of 20 to 50 using LASSO then fit an interpretable parametric model, applying post-double-selection when causal inference is the goal.
  • Elastic net for correlated code families (ICD-10 chapters, ATC drug classes) where the L1/L2 mixture retains the group rather than an arbitrary single member.
  • Nuisance model fitting inside double-ML or TMLE as one candidate in a super learner, where prediction quality — not causal interpretability — governs the nuisance fit.
  • Penalized Cox regression for time-to-event outcomes in linked registries when many candidate code predictors are available and OLS-based variable selection is infeasible.

When NOT to use

  • Small p where domain knowledge should select covariates: if a pre-specified set of 8 to 12 clinically motivated confounders is available, fit them in an unpenalized model; adding LASSO introduces bias without the high-dimensional benefit.
  • Primary causal effect estimation with naive SEs: never report a LASSO-selected and refitted coefficient as an unbiased causal estimate without post-double-selection or an equivalent causal ML approach.
  • Instruments or confounders chosen purely by predictive criteria: LASSO selects on outcome prediction; a code highly predictive of the outcome but connected to the exposure only through an unmeasured common cause is a collider — selecting it can induce M-bias or amplify unmeasured confounding. Data-adaptive selection cannot substitute for a causal DAG (see bias discussion in predictive-and-causal-ml-models-rwe).
  • When the complete covariate set must be retained for regulatory or scientific reasons: a pre-specified confirmatory analysis with a small number of adjudicated covariates should use unpenalized regression.
  • When LASSO selection is the deliverable but collinearity is severe: use elastic net or group lasso instead; pure LASSO with correlated predictors produces an arbitrary sparse representation of the correlation structure, not a scientifically meaningful feature set.

Citations

FOUNDATIONAL / METHODS
  1. [1]Tibshirani R. Regression Shrinkage and Selection via the Lasso. Journal of the Royal Statistical Society Series B. 1996;58(1):267-288.
  2. [2]Zou H, Hastie T. Regularization and variable selection via the elastic net. Journal of the Royal Statistical Society Series B. 2005;67(2):301-320.