Course 10, lesson 93 of 100, Adults
Backpropagation
Sharing out the blame
Like I’m 5
When a team loses a game, the coach works out who could have done what better, from the last pass back to the first. Backpropagation does this for every number inside an AI.
The big idea
Backpropagation computes the gradient of the loss with respect to every weight. First, a forward pass computes the output and loss. Then a backward pass applies the chain rule, layer by layer from the output back to the input, passing error signals along.
Frameworks like PyTorch do this automatically (automatic differentiation). Deep networks can suffer from vanishing or exploding gradients; residual connections, normalisation and careful initialisation keep the signal healthy so very deep models can train.
Examples
- Chain rule: If y = f(g(x)), then dy/dx = f′(g(x)) · g′(x).
- Autograd: loss.backward() fills in gradients for millions of weights.
- Residuals: Skip connections let gradients flow through very deep networks.
How it works
- Run a forward pass to get the output and the loss.
- Run a backward pass using the chain rule, from output to input.
- Use the gradients to update every weight.
Check your understanding
- Which maths rule powers backpropagation?
- Options: The chain rule; The rule of thirds; Pythagoras' theorem.
Answer: The chain rule. It composes derivatives through each layer. - What problem do residual connections help with?
- Options: Vanishing gradients in deep networks; Slow internet; Too much data.
Answer: Vanishing gradients in deep networks. Skip paths keep gradient signals strong through many layers.
Remember
Backprop uses the chain rule to send error signals backwards, giving every weight its gradient.
Talk about it
How would you share the credit and blame fairly in a group project?
Go deeper
Reverse-mode automatic differentiation computes all gradients in roughly the cost of a couple of forward passes, which is why training billion-parameter models is feasible.