The stochastic gradient descent algorithm (Stochastic Gradient Descent, SGD) is a simplification of standard gradient descent.
Instead of computing the exact gradient of the objective function
, at each iteration it uses the gradient of a single randomly selected sample
:
| (4.43) |
SGD is guaranteed to converge for convex functions defined on convex domains, but it is also commonly used in nonconvex settings, such as neural-network training.
A practical variant consists of using a small group of samples (mini-batch) for each update, reducing the noise relative to a single sample while maintaining good computational efficiency.
To simulate inertia in parameter updates, a term called momentum is introduced:
| (4.44) |
In addition to SGD with momentum, numerous variants have been designed to accelerate convergence and improve optimization stability. A non-exhaustive list includes:
Paolo medici