Backpropagation

How neural networks learn from their mistakes—from intuitive blame assignment and the Chain Rule to matrix backpropagation and advanced memory optimization.

28 minIntermediateCode Examples

The Core Thesis: When a deep neural network makes a wrong prediction at its final output layer, how does a weight buried 50 layers deep know whether it should move up or down? Backpropagation (short for "Backward Propagation of Errors") is the algorithm that solves this Credit Assignment Problem—walking backward from the final Loss layer by layer using the Chain Rule so every single weight learns its exact share of the blame in a single pass.

  1. Beginner Foundations: The Credit Assignment Problem & The Relay Race

Imagine a factory assembly line with three workers in a row: Worker 1 cuts the wood, Worker 2 assembles the chair, and Worker 3 paints it. At the end of the line, the quality inspector measures the final product and finds an error (The Loss LL).

Who is to blame, and how should each worker adjust their work? We cannot just yell at Worker 3! Instead, the inspector tells Worker 3 how much the paint contributed to the error, Worker 3 passes the remaining blame back to Worker 2, and Worker 2 passes the blame back to Worker 1. That two-way process is how every neural network trains:

1. The Forward Pass (Left to Right →)

Input data x\mathbf{x} flows forward through every layer to produce a prediction y^\hat{y} and measure the error L(y^,y)L(\hat{y}, y). Along the way, each layer saves (caches) its intermediate activations in memory.

Input x → Layer 1 → Layer 2 → Prediction → Loss L

2. The Backward Pass (Right to Left ←)

Starting at the Loss LL, error signals (gradients) flow backward from the output layer to the first layer, multiplying local slopes via the Chain Rule to compute ∂L∂w\dfrac{\partial L}{\partial w} for every weight.

dL/dW1 ← Layer 1 ← Layer 2 ← dL/dPrediction ← Loss L

Backpropagation vs. Gradient Descent (Do Not Confuse Them!)

Beginners often think Backpropagation and Gradient Descent are the same thing. They are two separate steps: Backpropagation is the algorithm that calculates the gradients (∇L\nabla L, the compass reading) using the Chain Rule. Gradient Descent (or Adam) is the optimizer that uses those gradients to update the weights (wnew=w−α∇Lw_{\text{new}} = w - \alpha \nabla L)!

  1. Beginner Walkthrough: Backpropagation Through One Neuron by Hand

Let us trace the exact numbers of a Forward Pass and Backward Pass through a single neuron with input x=2x = 2, weight w=3w = 3, bias b=1b = 1, ReLU activation a=ReLU(z)a = \text{ReLU}(z), and target y=10y = 10 using Squared Error Loss L=(a−y)2L = (a - y)^2:

z=wx+b⟶a=ReLU(z)⟶L=(a−y)2z = wx + b \quad \longrightarrow \quad a = \text{ReLU}(z) \quad \longrightarrow \quad L = (a - y)^2
Step A: Forward Pass (Compute & Save)
1. Linear sum: z=(3)(2)+1=7z = (3)(2) + 1 = 7
2. ReLU switch: a=max⁡(0,7)=7a = \max(0, 7) = 7
3. Error (Loss): L=(7−10)2=(−3)2=9L = (7 - 10)^2 = (-3)^2 = 9
Step B: Backward Pass (Right to Left)
1. Loss slope: ∂L∂a=2(a−y)=2(7−10)=−6\dfrac{\partial L}{\partial a} = 2(a - y) = 2(7 - 10) = -6
2. Through ReLU (z>0  ⟹  slope 1z \gt 0 \implies \text{slope } 1): ∂L∂z=(−6)(1)=−6\dfrac{\partial L}{\partial z} = (-6)(1) = -6
3. Weight & Bias gradients: ∂L∂w=∂L∂z⋅x=(−6)(2)=−12\dfrac{\partial L}{\partial w} = \dfrac{\partial L}{\partial z} \cdot x = (-6)(2) = -12, and ∂L∂b=−6\dfrac{\partial L}{\partial b} = -6
⚡ Knowledge Check

In the neuron example above (z=wx+bz = wx + b), why did we need to know the forward input value x=2x = 2 during the Backward Pass when calculating ∂L∂w\dfrac{\partial L}{\partial w}?

A) Because the local derivative dz/dw equals x, so dL/dw = (dL/dz) * x▼
✓ Correct!Because z=wx+bz = wx + b, the partial derivative ∂z∂w\dfrac{\partial z}{\partial w} is literally the input xx! That is why neural networks must store forward activations in GPU RAM until backpropagation finishes.
B) Because we are updating the input data x instead of w▼
✕ Incorrect.Input data xx is fixed. We use xx purely because the sensitivity of a multiplier gate (w⋅xw \cdot x) with respect to ww is the other multiplier xx.

  1. Intermediate: The Four Fundamental Equations of Backpropagation

Now let us scale up from a single neuron to a Multilayer Neural Network with layers l=1,2,…,Ll = 1, 2, \dots, L. At every layer ll, we define an Error Signal Vector δ(l)=∂L∂z(l)\boldsymbol{\delta}^{(l)} = \dfrac{\partial \mathcal{L}}{\partial \mathbf{z}^{(l)}}, which represents the gradient of the Loss with respect to that layer's pre-activation logits z(l)\mathbf{z}^{(l)}.

All of backpropagation across any deep network boils down to four famous equations (where ⊙\odot is element-wise multiplication):

Equation 1: Output Layer Error (δ(L)\boldsymbol{\delta}^{(L)})

Computes the error signal at the final layer LL. When using Softmax + Cross-Entropy (or Sigmoid + BCE), the derivative simplifies miraculously to just Prediction minus Truth!

δ(L)=∂L∂z(L)=y^−y\boldsymbol{\delta}^{(L)} = \dfrac{\partial \mathcal{L}}{\partial \mathbf{z}^{(L)}} = \hat{\mathbf{y}} - \mathbf{y}
Equation 2: Pass Error to Previous Layer (δ(l)\boldsymbol{\delta}^{(l)})

Routes the next layer's error δ(l+1)\boldsymbol{\delta}^{(l+1)} backward through the transposed weight matrix (W(l+1))T(\mathbf{W}^{(l+1)})^T and multiplies element-wise by the activation slope σ′(z(l))\sigma^\prime(\mathbf{z}^{(l)}):

δ(l)=(δ(l+1)(W(l+1))T)⊙σ′(z(l))\boldsymbol{\delta}^{(l)} = \left(\boldsymbol{\delta}^{(l+1)} \left(\mathbf{W}^{(l+1)}\right)^T\right) \odot \sigma^\prime\left(\mathbf{z}^{(l)}\right)
Equation 3: Weight Matrix Gradient (∂L∂W(l)\dfrac{\partial \mathcal{L}}{\partial \mathbf{W}^{(l)}})

Once layer ll knows its error signal δ(l)\boldsymbol{\delta}^{(l)}, the gradient for its weight matrix W(l)\mathbf{W}^{(l)} is simply the outer product (matrix multiply) of its incoming activations and its error signal:

∂L∂W(l)=(a(l−1))Tδ(l)\dfrac{\partial \mathcal{L}}{\partial \mathbf{W}^{(l)}} = \left(\mathbf{a}^{(l-1)}\right)^T \boldsymbol{\delta}^{(l)}
Equation 4: Bias Vector Gradient (∂L∂b(l)\dfrac{\partial \mathcal{L}}{\partial \mathbf{b}^{(l)}})

Because the bias b(l)\mathbf{b}^{(l)} is added directly to z(l)=a(l−1)W(l)+b(l)\mathbf{z}^{(l)} = \mathbf{a}^{(l-1)}\mathbf{W}^{(l)} + \mathbf{b}^{(l)} with a local slope of 11, the bias gradient is just the error signal summed across the batch!

∂L∂b(l)=∑batchδ(l)\dfrac{\partial \mathcal{L}}{\partial \mathbf{b}^{(l)}} = \sum_{\text{batch}} \boldsymbol{\delta}^{(l)}

Why Does Equation 2 Multiply by the Transposed Weight Matrix WT\mathbf{W}^T?

During the Forward Pass, matrix W\mathbf{W} of shape (din,dout)(d_{\text{in}}, d_{\text{out}}) projects signals forward from dind_{\text{in}} neurons to doutd_{\text{out}} neurons. During the Backward Pass, we need to send error signals in the exact opposite direction—from doutd_{\text{out}} neurons back to dind_{\text{in}} neurons—which requires multiplying by WT\mathbf{W}^T of shape (dout,din)(d_{\text{out}}, d_{\text{in}})!

  1. Intermediate: Why Backpropagation is 1,000,000x Faster (Dynamic Programming)

Why was the popularization of Backpropagation by Rumelhart, Hinton, and Williams in 1986 such a historic breakthrough? Consider how long it would take to train a small 1,000,0001{,}000{,}000-parameter network without Backpropagation:

MethodHow It Computes All NN GradientsCompute Cost per StepTime for 1M Weights (if 1 pass = 1 ms)
Finite Differences (Forward Perturbation)Nudges 1 weight at a time (wi+ϵw_i + \epsilon) and runs a full forward pass per weight.O(N)\mathcal{O}(N) Forward Passes1,000 seconds (~17 mins) per step!
Backpropagation (Reverse-Mode Autodiff)Reuses cached intermediate derivatives from right to left in one sweep.O(1)\mathcal{O}(1) Backward Pass (≈2×\approx 2\times Forward)0.002 seconds (2 ms) for ALL 1M weights!
⚡ Knowledge Check

In a hidden layer, the incoming activation batch A(l−1)\mathbf{A}^{(l-1)} has shape (32,128)(32, 128) and the upstream error signal δ(l)\boldsymbol{\delta}^{(l)} has shape (32,64)(32, 64). What is the shape of the weight gradient ∂L∂W(l)=(A(l−1))Tδ(l)\dfrac{\partial \mathcal{L}}{\partial \mathbf{W}^{(l)}} = \left(\mathbf{A}^{(l-1)}\right)^T \boldsymbol{\delta}^{(l)}?

A) (128, 64) — matching the exact shape of the weight matrix W!▼
✓ Correct!Transposing A(l−1)\mathbf{A}^{(l-1)} gives (128,32)(128, 32). Multiplying (128,32)×(32,64)(128, 32) \times (32, 64) sums over the 3232 batch samples and yields (128,64)(128, 64), which matches W(l)\mathbf{W}^{(l)}'s shape (128,64)(128, 64)!
B) (32, 32)▼
✕ Incorrect.Remember the Golden Shape Rule: a weight matrix's gradient ∂L∂W\dfrac{\partial \mathcal{L}}{\partial \mathbf{W}} must always have the exact same dimensions (din,dout)=(128,64)(d_{\text{in}}, d_{\text{out}}) = (128, 64) as W\mathbf{W} itself.

  1. Advanced Production Backprop: Activation Memory, Checkpointing & Gradient Clipping

When training Large Language Models (like Llama-3 or GPT-4) on GPU clusters, senior AI engineers face three critical systems-level challenges directly caused by Backpropagation:

1. Why Training Uses 10x More VRAM Than Inference

Because Equation 3 (ATδ\mathbf{A}^T \boldsymbol{\delta}) requires every layer's forward activation A(l)\mathbf{A}^{(l)} during the backward pass, PyTorch must keep all intermediate activations from all layers in GPU VRAM! During inference, activations are discarded immediately after each layer.

Training VRAM = Weights + Gradients + Optimizer + All Activations!

2. Activation (Gradient) Checkpointing

How do we train a 70B LLM without running out of GPU memory? Activation Checkpointing discards most intermediate activations during the forward pass and recomputes them on the fly during the backward pass—trading ≈25%\approx 25\% extra compute to slash activation VRAM by 70%+70\%+!

torch.utils.checkpoint.checkpoint(layer_block, x)

  1. Stabilizing Backprop in Deep Networks: Gradient Clipping & Residual Highways

When training Transformers or RNNs, an unlucky batch can cause the backward product of matrices to spike (Exploding Gradients). Before calling optimizer.step(), production training loops cap the global gradient norm to 1.01.0 using torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0) so a single bad batch never destroys the model's learned weights.

⚡ Knowledge Check

Your 7-Billion parameter LLM fits easily on a single 24GB GPU during inference (with torch.no_grad():), but immediately crashes with CUDA out of memory on the exact same GPU when you try to run training with batch size 88. Why?

A) Backprop must cache forward activations for all layers, plus store gradients and optimizer states▼
✓ Correct!During inference, layer 11's activations are overwritten as soon as layer 22 finishes. During training, every layer's activations must be kept in VRAM until the backward pass reaches that layer. Enabling Activation Checkpointing solves this!
B) Because training duplicates the input dataset 100 times in VRAM▼
✕ Incorrect.Only one mini-batch of data is loaded into VRAM at a time; the memory spike comes from cached intermediate layer activations, parameter gradients, and Adam optimizer momentum buffers.

  1. Visual Explanation: The Complete Forward & Backward Pass Loop

Look at how signals travel forward through the network to cache activations and compute the Loss, and how error signals δ\boldsymbol{\delta} travel backward through transposed weights to compute parameter gradients:

  1. Python Implementation: Manual Matrix Backprop vs. PyTorch Autograd Verification

The ultimate test of understanding Backpropagation is implementing the 4 matrix equations yourself and verifying that your manual gradients match PyTorch's loss.backward() to 77 decimal places:

backprop_from_scratch.pyPython 3.11+ · PyTorch Verification
import torch # 1. Create Batch X (4 samples, 3 features) and Binary Targets Y (4, 1)torch.manual_seed(7)X = torch.randn(4, 3)Y = torch.tensor([[1.0], [0.0], [1.0], [0.0]])batchSize = X.shape[0] # Initialize 2-Layer Network Weights (Requires Grad for PyTorch Check)W1 = torch.randn(3, 5, requires_grad=True)b1 = torch.zeros(1, 5, requires_grad=True)W2 = torch.randn(5, 1, requires_grad=True)b2 = torch.zeros(1, 1, requires_grad=True) # 2. FORWARD PASS (Compute Predictions & Cache Z1, A1)Z1 = X @ W1 + b1A1 = torch.relu(Z1)Z2 = A1 @ W2 + b2yHat = torch.sigmoid(Z2)loss = -torch.mean(Y * torch.log(yHat) + (1 - Y) * torch.log(1 - yHat)) # 3. MANUAL BACKPROPAGATION (The 4 Fundamental Equations!)with torch.no_grad():    delta2 = (yHat - Y) / batchSize                   # Eq 1: Output Error (4, 1)    dW2Manual = A1.T @ delta2                         # Eq 3: Layer 2 Weight Grad (5, 1)    db2Manual = torch.sum(delta2, dim=0, keepdim=True)# Eq 4: Layer 2 Bias Grad (1, 1)     delta1 = (delta2 @ W2.T) * (Z1 > 0).float()       # Eq 2: Hidden Error (4, 5)    dW1Manual = X.T @ delta1                          # Eq 3: Layer 1 Weight Grad (3, 5)    db1Manual = torch.sum(delta1, dim=0, keepdim=True)# Eq 4: Layer 1 Bias Grad (1, 5) # 4. Run PyTorch Autograd and Verify Exact Match!loss.backward()print("Max diff in dW1:", torch.max(torch.abs(dW1Manual - W1.grad)).item())  # ~0.0print("Max diff in dW2:", torch.max(torch.abs(dW2Manual - W2.grad)).item())  # ~0.0

Pro Tip (Why Equation 1 Is So Clean): Look at delta2 = (yHat - Y) / batchSize in the code above! When you pair Sigmoid with Binary Cross-Entropy (or Softmax with Multi-Class Cross-Entropy), the messy derivatives of the logarithm and the exponential cancel each other out completely, leaving just y^−y\hat{\mathbf{y}} - \mathbf{y}!

Key Points

✓Backpropagation is the reverse-mode automatic differentiation algorithm that computes the gradient ∇L\nabla \mathcal{L} for every weight in a neural network in a single backward pass.
✓Backpropagation only computes the gradients; an optimizer like SGD or AdamW then uses those gradients to update the weights[cite: 10].
✓The four equations of matrix backpropagation compute: (1) output error δ(L)\boldsymbol{\delta}^{(L)}, (2) hidden error δ(l)=(δ(l+1)(W(l+1))T)⊙σ′(z(l))\boldsymbol{\delta}^{(l)} = (\boldsymbol{\delta}^{(l+1)}(\mathbf{W}^{(l+1)})^T) \odot \sigma^\prime(\mathbf{z}^{(l)}), (3) weight gradient (A(l−1))Tδ(l)(\mathbf{A}^{(l-1)})^T \boldsymbol{\delta}^{(l)}, and (4) bias gradient ∑δ(l)\sum \boldsymbol{\delta}^{(l)}.
✓By caching and reusing intermediate derivatives from right to left (Dynamic Programming), Backpropagation computes NN parameter gradients in O(1)\mathcal{O}(1) passes instead of O(N)\mathcal{O}(N) passes.
✓Because computing ∂L∂W(l)\dfrac{\partial \mathcal{L}}{\partial \mathbf{W}^{(l)}} requires the forward activation A(l−1)\mathbf{A}^{(l-1)}, training a deep model consumes far more GPU memory than inference unless Activation Checkpointing is enabled.

Common Mistakes

✕ Confusing Backpropagation with Gradient Descent.

Backpropagation is the algorithm that calculates ∂L∂W\dfrac{\partial \mathcal{L}}{\partial \mathbf{W}} via the Chain Rule (called by loss.backward()). Gradient Descent / Adam is the optimization rule that steps downhill using those gradients (called by optimizer.step())[cite: 10].

✕ Using matrix multiplication (@) instead of element-wise multiplication (*) for activation derivatives.

In Equation 2, (δ(l+1)(W(l+1))T)\left(\boldsymbol{\delta}^{(l+1)} (\mathbf{W}^{(l+1)})^T\right) is a matrix multiplication, but multiplying by the activation slope σ′(z(l))\sigma^\prime(\mathbf{z}^{(l)}) is strictly element-wise (⊙\odot) because each neuron gates its own signal.

✕ Forgetting to divide the output gradient by batch_size when using mean loss.

If your forward loss averages error over BB batch samples (1B∑Li\frac{1}{B}\sum \mathcal{L}_i), your backward output error δ(L)\boldsymbol{\delta}^{(L)} must also be divided by BB so gradient magnitudes do not scale up when you change the batch size!

✕ Modifying forward activations in-place before backpropagation runs.

Overwriting cached forward tensors with in-place operations (like a += 1) corrupts the values needed for Equation 3 (ATδ\mathbf{A}^T \boldsymbol{\delta}), causing PyTorch's Autograd engine to throw a RuntimeError.

The Big Picture

Forward Perturbation (Changing 1 Weight at a Time)

1 Billion Weights → Requires 1 Billion Forward Passes per Step → Impossible to Train Deep AI

Backpropagation (Reverse-Mode Chain Rule)

1 Forward Pass (Cache Activations) → 1 Backward Sweep (Pass Error δ Right-to-Left) → All 1B Gradients at Once!

The important conceptual shift is realizing that Backpropagation is what made modern Deep Learning computationally possible. By saving activations during the forward pass and passing error signals backward through transposed weight matrices, a GPU can compute the exact slope for billions of parameters in roughly the same time it takes to make a single prediction.

Remember: Signals flow forward through W\mathbf{W} to make predictions, and blame flows backward through WT\mathbf{W}^T via Backpropagation so every layer learns from the final mistake.