All articles
Deep Learning·Fundamentals

Backpropagation, Step by Step

How the chain rule turns a network's forward pass into gradients for every weight, and why the "backward" direction is what makes training tractable.

·2 min read·
2 views
deep-learningmathneural-networks
Notes / Deep Learning

Backpropagation, Step by Step

#deep-learning#math#neural-networks

Backpropagation, Step by Step

Backpropagation is the algorithm that computes how much each weight in a neural network contributed to the final error, so gradient descent knows which direction to move each one. It is nothing more than the chain rule, applied systematically through a computation graph.

Forward Pass

Consider a tiny two-layer network: an input xx, a hidden unit h=σ(wx+b)h = \sigma(wx + b), and an output y^=vh\hat{y} = vh, trained with squared error loss L=(y^−y)2\mathcal{L} = (\hat{y} - y)^2.

The forward pass computes each quantity left to right: x→h→y^→Lx \to h \to \hat{y} \to \mathcal{L}, caching every intermediate value along the way. Those cached values are what backpropagation reuses — recomputing them during the backward pass would waste work.

Backward Pass: The Chain Rule

To update ww, we need ∂L/∂w\partial \mathcal{L} / \partial w. The chain rule decomposes it into a product of local derivatives, one per step in the forward pass:

∂L∂w=∂L∂y^⋅∂y^∂h⋅∂h∂w\frac{\partial \mathcal{L}}{\partial w} = \frac{\partial \mathcal{L}}{\partial \hat{y}} \cdot \frac{\partial \hat{y}}{\partial h} \cdot \frac{\partial h}{\partial w}

Each factor is easy to compute in isolation:

TermValue
∂L/∂y^\partial \mathcal{L} / \partial \hat{y}2(y^−y)2(\hat{y} - y)
∂y^/∂h\partial \hat{y} / \partial hvv
∂h/∂w\partial h / \partial wσ′(wx+b)⋅x\sigma'(wx+b) \cdot x

Multiplying them together gives the gradient with respect to ww, without ever writing out the full expression for L\mathcal{L} in terms of ww.

Why "Back"-propagation

The key insight is order: computing ∂L/∂y^\partial \mathcal{L}/\partial \hat{y} first and reusing it for every downstream derivative is far cheaper than recomputing the chain from scratch for every weight. In a network with millions of parameters sharing intermediate layers, this reuse is what makes training tractable — the cost of backpropagation is roughly the same as a single forward pass, regardless of how many parameters you have.

A Toy Implementation

def backward(x, y, w, b, v):
    h = sigmoid(w * x + b)
    y_hat = v * h
    dL_dyhat = 2 * (y_hat - y)
    dyhat_dh = v
    dh_dw = sigmoid_derivative(w * x + b) * x
 
    dL_dw = dL_dyhat * dyhat_dh * dh_dw
    dL_dv = dL_dyhat * h
    return dL_dw, dL_dv

Autodiff frameworks like PyTorch and JAX generalize exactly this pattern to arbitrary computation graphs: every operation records its local derivative, and a single backward traversal multiplies them along every path from the loss back to each parameter.

Newsletter

New articles in your inbox

An email when I publish something new. No spam, unsubscribe anytime.

Double opt-in. See the privacy policy for how your email is handled.

Continue exploring

Browse more technical notes and experiments.

All articles