Adam and similar methods fail to converge in some deep learning tasks.
problem Adam and similar methods fail to converge to optimal solutions in some deep learning tasks.
method Analysis of gradient updates scaled by square roots of exponential moving averages of squared past gradients. Proposed new variants of Adam with `long-term memory'.
result New variants of Adam algorithms fix convergence issues and often lead to improved empirical performance.