Adam (KB14) (Adaptive Moment Estimation) is one of the most widely used optimization algorithms in deep learning, owing to its ability to combine the advantages of AdaGrad and RMSProp.
AdaGrad assigns a specific learning rate to each parameter, making it effective in the presence of sparse gradients. RMSProp, on the other hand, adapts the learning rate according to the gradient magnitude, making it suitable for online and nonstationary scenarios.
Adam extends these approaches by introducing an adaptive estimate of both the first-order moment (the mean of the gradients) and the second-order moment (the variance of the gradients).
In particular, for each model
, Adam maintains two variables:
These estimates are updated iteratively as follows:
| (4.42) |
| (4.43) |
Since and
are initialized to zero, the first iterations are underestimated. To correct this bias, the following are computed:
| (4.44) |
The parameters are then updated according to:
| (4.45) |
Adam is particularly effective in scenarios with noisy data, sparse gradients, or nonstationary objective functions. Thanks to its robustness and ease of use, it is often the default choice for optimizing deep neural networks.
Paolo medici