The stochastic gradient descent algorithm (Stochastic Gradient Descent, SGD) is a simplification of classical gradient descent.
Instead of computing the exact gradient of the objective function
, at each iteration the gradient of a single randomly selected sample
is used:
| (4.40) |
SGD is guaranteed to converge for convex functions defined on convex domains, but it is also commonly used in nonconvex settings, such as neural-network training.
A practical variant consists of using a small group of samples (mini-batch) for each update, reducing the noise relative to a single sample while maintaining good computational efficiency.
To simulate inertia in the parameter updates, a term called momentum is introduced:
| (4.41) |
In addition to SGD with momentum, numerous variants have been designed to accelerate convergence and improve optimization stability. A non-exhaustive list includes:
Paolo medici