Course 10, lesson 92 of 100, Adults
Loss and gradient descent
How models get less wrong
Like I’m 5
Imagine you're on a foggy hill and want to reach the bottom. You feel which way slopes down and take a small step. Repeat, and you get there. That's how AI learns.
The big idea
A loss function measures how wrong the model is: mean squared error for numbers, cross-entropy for categories. Training means finding weights that make the loss small.
The gradient is the direction in which the loss increases fastest, so we step the opposite way: w ← w − η∇L, where η is the learning rate. Too big a step overshoots; too small crawls. In practice we use mini-batches (stochastic gradient descent) and optimisers like Adam.
Examples
- Squared error: Predicted 7, true 10: loss (7 − 10)² = 9.
- Learning rate: 0.1 might bounce around; 0.0001 might take forever.
- Mini-batches: Update after every 64 examples instead of the whole dataset.
How it works
- Compute the loss on a batch of examples.
- Compute the gradient of the loss for every weight.
- Step each weight a little in the downhill direction, and repeat.
Check your understanding
- Gradient descent moves weights in which direction?
- Options: Opposite to the gradient, downhill; Along the gradient, uphill; In a random direction.
Answer: Opposite to the gradient, downhill. The gradient points uphill, so we step the other way. - What happens if the learning rate is far too large?
- Options: Training can overshoot and diverge; Training becomes perfect; Nothing changes.
Answer: Training can overshoot and diverge. Huge steps jump past the minimum and can blow up.
Remember
Loss measures error; gradient descent repeatedly steps weights downhill to reduce it.
Talk about it
When have you improved something by making small adjustments and checking?
Go deeper
Adam combines momentum with per-parameter adaptive step sizes. Learning-rate schedules with warm-up and decay, plus weight decay, are standard in large-model training.