LLMpediaThe first transparent, open encyclopedia generated by LLMs

Lasso (statistics)

Note: This article was automatically generated by a large language model (LLM) from purely parametric knowledge (no retrieval). It may contain inaccuracies or hallucinations. This encyclopedia is part of a research project currently under review.
Article Genealogy
Parent: Frank E. Harrell Jr. Hop 5 terminal

This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.

Lasso (statistics)
NameLasso
ClassificationRegularization method
Introduced1996
InventorRobert Tibshirani
RelatedRidge regression, Least absolute deviations, Elastic Net

Lasso (statistics) The Lasso is a regression method that performs variable selection and regularization by imposing an L1 penalty on coefficients to produce sparse solutions. Originally proposed by Robert Tibshirani in 1996, the method connects to earlier work by Trevor Hastie, Jerome Friedman, Bradley Efron, and ideas from Geoffrey Hinton's sparse representations, while influencing developments involving Hastie–Tibshirani–Friedman frameworks and tools like Scikit-learn, R Project, MATLAB toolboxes.

Introduction

The Lasso was introduced to address high-dimensional estimation problems encountered in fields such as Stanford University-affiliated genomics projects, Harvard University-based econometrics studies, and large-scale surveys from organizations like National Institutes of Health and European Bioinformatics Institute. It stands alongside predecessors like Ridge regression and successors like Elastic Net as a cornerstone in modern statistical learning curricula influenced by texts from Christopher Bishop, Trevor Hastie, and Robert Tibshirani.

Mathematical formulation

The canonical Lasso estimator minimizes a penalized least-squares objective: given design matrix X and response y, coefficients β solve argmin_{β} ||y - Xβ||_2^2 + λ||β||_1. This formulation is rooted in convex analysis developed in work by John von Neumann-inspired linear programming and later convex optimization theory by Yurii Nesterov and Michel Goemans; the penalty parameter λ connects to model selection criteria such as Akaike Information Criterion and Bayesian Information Criterion in practice. Equivalent characterizations use subgradient conditions and KKT optimality as studied in papers by David Donoho, Emmanuel Candès, and Iain Johnstone.

Computational algorithms

Algorithms for solving the Lasso include coordinate descent, least-angle regression (LARS), proximal gradient methods, and interior-point methods. The LARS algorithm, introduced by Bradley Efron with collaborators, computes an entire solution path efficiently, while coordinate descent implementations are popular in Scikit-learn and glmnet by Jerome Friedman and colleagues, leveraging warm starts and active-set strategies. Proximal algorithms and accelerated methods trace to work by Yurii Nesterov and packages in CVX and MOSEK.

Statistical properties and theory

Theoretical analysis covers consistency, sparsity recovery, oracle inequalities, and asymptotic distribution. Notable results include sparsistency conditions like the irrepresentable condition studied by Peter Bühlmann and Sara van de Geer, minimax optimality results related to work by David Donoho and Iain Johnstone, and bounds under restricted isometry properties developed in compressive sensing by Emmanuel Candès and Terence Tao. Debiased or desparsified Lasso estimators for inference were advanced by researchers such as Mark van de Geer and Peter Bühlmann facilitating confidence intervals and hypothesis tests in high-dimensional regimes.

Extensions and variants

Numerous variants extend the Lasso: the Elastic Net combines L1 and L2 penalties (credited to Hui Zou and Trevor Hastie), Adaptive Lasso uses data-driven weights (proposed by Hui Zou), Group Lasso imposes group sparsity (developed by Mauro Yuan and colleagues), Fused Lasso encourages total variation regularization (introduced by Robert Tibshirani and James Taylor), and Graphical Lasso estimates sparse precision matrices (popularized by Jerome Friedman and Trevor Hastie in graphical models research). Other adaptations include SCAD and MCP penalties proposed by Jianqing Fan and Runze Li for nonconvex regularization, and square-root Lasso by Mathias van de Geer and others for heteroscedastic settings.

Applications

The Lasso has been applied across genomics in studies at Broad Institute, neuroimaging projects at Massachusetts General Hospital, finance research at Goldman Sachs and BlackRock, natural language processing in labs at Google and OpenAI, and environmental modeling in collaborations with NASA and European Space Agency. It underpins feature selection in competitions like Netflix Prize and informs pathway analysis in studies funded by National Institutes of Health and Wellcome Trust.

Practical considerations and implementation

Practical use involves standardizing predictors, selecting λ via cross-validation (K-fold schemes popularized in software like Scikit-learn and glmnet), and interpreting sparse solutions with caution when predictors are highly collinear (a concern addressed by Elastic Net and principal components methods from Karl Pearson and Harold Hotelling). Software implementations exist in R Project packages such as glmnet and interfaces in Python libraries like Scikit-learn, with efficient compiled backends from contributors affiliated with Stanford University and University of California, Berkeley.

Category:Statistical learning theory