All articles
Machine Learning·Optimization

Gradient Descent, Intuitively

Why almost every model learns by walking downhill: the loss landscape, the role of the learning rate, and how mini-batches change the walk.

·2 min read·
0 views
beginnersmachine-learningoptimization
Notes / Machine Learning

Gradient Descent, Intuitively

#beginners#machine-learning#optimization

Gradient Descent, Intuitively

Almost every model you train — linear regression, a CNN, a transformer — learns the same way: it nudges its parameters a little at a time to make its predictions less wrong. That nudging process is gradient descent.

The Loss Landscape

Picture every possible setting of a model's parameters as a point on a landscape, and the model's error at that point as the landscape's height. Training is walking downhill until you reach a valley — a set of parameters where the error is low.

θt+1=θt−η∇θL(θt)\theta_{t+1} = \theta_t - \eta \nabla_\theta \mathcal{L}(\theta_t)

Here θ\theta is the parameter vector, L\mathcal{L} is the loss function, ∇θL\nabla_\theta \mathcal{L} is the gradient (the direction of steepest ascent), and η\eta is the learning rate. Subtracting the gradient moves you downhill; the learning rate controls the size of each step.

Why the Learning Rate Matters

Learning rateWhat happens
Too smallTraining crawls; may get stuck in a shallow dip long before it matters
Too largeSteps overshoot the valley and the loss oscillates or diverges
Well-tunedSteady, fast descent toward a good minimum

In practice, the rate is rarely constant. Schedulers shrink it over time, and adaptive optimizers like Adam rescale it per-parameter based on the recent history of gradients.

Batches: Full, Stochastic, and Mini

Computing the exact gradient requires averaging the loss over the entire dataset, which is expensive. Stochastic gradient descent (SGD) estimates it from a single example instead — noisy, but cheap. Mini-batch gradient descent splits the difference, averaging over a small batch (say, 32 or 256 examples) at each step. That noise is not purely a nuisance: it helps the optimizer escape shallow local minima and saddle points that a perfectly smooth gradient would settle into.

A Minimal Implementation

def gradient_descent(grad_fn, theta, lr=0.01, steps=1000):
    for _ in range(steps):
        grad = grad_fn(theta)
        theta = theta - lr * grad
    return theta

grad_fn computes ∇θL\nabla_\theta \mathcal{L} at the current parameters — for a neural network, this is exactly what backpropagation produces. Everything else in modern deep learning optimizers (momentum, Adam, weight decay) is a refinement of this same loop: look at the slope, take a step, repeat.

What Is AI?

A beginner-friendly walkthrough of what artificial intelligence actually is, how it differs from ordinary software, and where AI, machine learning, deep learning, and LLMs fit together.

How LLMs Work

A plain-language walkthrough of what actually happens between typing a prompt and a large language model typing back an answer: tokens, embeddings, attention, and why they sometimes make things up.

Word Embeddings and Why They Work

The distributional hypothesis, the skip-gram objective, and why vector arithmetic on words like king minus man plus woman approximates queen.

Newsletter

New articles in your inbox

An email when I publish something new. No spam, unsubscribe anytime.

Double opt-in. See the privacy policy for how your email is handled.

Continue exploring

Browse more technical notes and experiments.

All articles