Watch five optimizers race across a 2D loss landscape.
Each terrain has an analytic gradient. Every frame, the five optimizers take one step from a shared, draggable start, leaving fading trails. SGD follows the raw gradient; Momentum and NAG accumulate velocity; RMSProp and Adam adapt the step per coordinate. Watch Adam tunnel through ravines and escape saddles where plain SGD stalls.
The geometry an optimizer sees — narrow valleys, flat saddles, jagged local minima — decides whether training converges fast, slowly, or not at all. Adaptive methods trade a little theory for huge practical robustness, which is why Adam dominates deep learning.