Stochastic Gradient Descent

The stochastic gradient descent algorithm (Stochastic Gradient Descent, SGD) is a simplification of classical gradient descent. Instead of computing the exact gradient of the objective function $S(\boldsymbol\beta)$, at each iteration the gradient of a single randomly selected sample $\ell_i$ is used:

\begin{displaymath}
\boldsymbol\beta_{t+1} = \boldsymbol\beta_{t} - \gamma_t \nabla \ell_i(\boldsymbol\beta_t)
\end{displaymath} (4.40)

SGD is guaranteed to converge for convex functions defined on convex domains, but it is also commonly used in nonconvex settings, such as neural-network training.

A practical variant consists of using a small group of samples (mini-batch) for each update, reducing the noise relative to a single sample while maintaining good computational efficiency.

To simulate inertia in the parameter updates, a term $\alpha$ called momentum is introduced:

\begin{displaymath}
\boldsymbol\delta_{t} = - \gamma_t \nabla S(\boldsymbol\beta_t) + \alpha \boldsymbol\delta_{t-1}
\end{displaymath} (4.41)

where $\alpha$ is usually a small value (e.g., 0.05). Momentum is one of the simplest and most effective modifications to SGD, helping to overcome oscillations and slowdowns in uninformative directions.

In addition to SGD with momentum, numerous variants have been designed to accelerate convergence and improve optimization stability. A non-exhaustive list includes:

A comparative overview of these algorithms is available in (Rud16).

Paolo medici
2026-10-01