A gentle look at gradient descent

2026-10-01

Gradient descent is the workhorse behind most of modern machine learning. The idea is simple: compute how the loss changes when each parameter changes, then take a small step in the opposite direction.

The learning rate decides the size of that step. Too large and the loss jumps around; too small and training takes forever. In practice a schedule that starts larger and slowly decays works well.

Variants such as momentum and Adam keep a running memory of previous gradients, which smooths the path and usually converges faster on noisy problems.