Gradient Descent, Intuitively
Why almost every model learns by walking downhill: the loss landscape, the role of the learning rate, and how mini-batches change the walk.
Gradient Descent, Intuitively
Gradient Descent, Intuitively
Almost every model you train — linear regression, a CNN, a transformer — learns the same way: it nudges its parameters a little at a time to make its predictions less wrong. That nudging process is gradient descent.
The Loss Landscape
Picture every possible setting of a model's parameters as a point on a landscape, and the model's error at that point as the landscape's height. Training is walking downhill until you reach a valley — a set of parameters where the error is low.
Here is the parameter vector, is the loss function, is the gradient (the direction of steepest ascent), and is the learning rate. Subtracting the gradient moves you downhill; the learning rate controls the size of each step.
Why the Learning Rate Matters
| Learning rate | What happens |
|---|---|
| Too small | Training crawls; may get stuck in a shallow dip long before it matters |
| Too large | Steps overshoot the valley and the loss oscillates or diverges |
| Well-tuned | Steady, fast descent toward a good minimum |
In practice, the rate is rarely constant. Schedulers shrink it over time, and adaptive optimizers like Adam rescale it per-parameter based on the recent history of gradients.
Batches: Full, Stochastic, and Mini
Computing the exact gradient requires averaging the loss over the entire dataset, which is expensive. Stochastic gradient descent (SGD) estimates it from a single example instead — noisy, but cheap. Mini-batch gradient descent splits the difference, averaging over a small batch (say, 32 or 256 examples) at each step. That noise is not purely a nuisance: it helps the optimizer escape shallow local minima and saddle points that a perfectly smooth gradient would settle into.
A Minimal Implementation
def gradient_descent(grad_fn, theta, lr=0.01, steps=1000):
for _ in range(steps):
grad = grad_fn(theta)
theta = theta - lr * grad
return thetagrad_fn computes at the current parameters — for a neural network, this is exactly what backpropagation produces. Everything else in modern deep learning optimizers (momentum, Adam, weight decay) is a refinement of this same loop: look at the slope, take a step, repeat.
Related notes
What Is AI?
A beginner-friendly walkthrough of what artificial intelligence actually is, how it differs from ordinary software, and where AI, machine learning, deep learning, and LLMs fit together.
How LLMs Work
A plain-language walkthrough of what actually happens between typing a prompt and a large language model typing back an answer: tokens, embeddings, attention, and why they sometimes make things up.
Word Embeddings and Why They Work
The distributional hypothesis, the skip-gram objective, and why vector arithmetic on words like king minus man plus woman approximates queen.
Newsletter
New articles in your inbox
An email when I publish something new. No spam, unsubscribe anytime.
Double opt-in. See the privacy policy for how your email is handled.