L1 and L2 Regularization

L1 and L2 regularization consists in adding an additional term to the cost function that penalizes certain configurations. For example, regularizing the cost function

\begin{displaymath}
S(\boldsymbol\beta, \mathbf{X}) = - \sum_i \log P (Y = y_i \vert \mathbf{x}_i ; \boldsymbol\beta)
\end{displaymath} (5.124)

means adding a term that is a function only of $\boldsymbol\beta$, thereby obtaining a new cost function of the form
\begin{displaymath}
E(\boldsymbol\beta, \mathbf{X}) = S(\boldsymbol\beta, \mathbf{X}) + \lambda R(\boldsymbol\beta)
\end{displaymath} (5.125)

where $R(\boldsymbol\beta)$ is a regularization function.

A widely used regularization function is

\begin{displaymath}
R(\boldsymbol\beta) = \left( \sum_j \vert \beta_j \vert ^ p \right)^{1/p}
\end{displaymath} (5.126)

Common values for $p$ are $1$ or $2$ (hence the names L1 and L2 regularization). When $p=2$, it may also be referred to in the literature as weight decay. These regularization functions therefore penalize parameters with excessively large values.



Paolo medici
2026-10-01