Course 10, lesson 92 of 100, Adults

Loss and gradient descent

How models get less wrong

Like I’m 5

Imagine you're on a foggy hill and want to reach the bottom. You feel which way slopes down and take a small step. Repeat, and you get there. That's how AI learns.

The big idea

A loss function measures how wrong the model is: mean squared error for numbers, cross-entropy for categories. Training means finding weights that make the loss small.

The gradient is the direction in which the loss increases fastest, so we step the opposite way: w ← w − η∇L, where η is the learning rate. Too big a step overshoots; too small crawls. In practice we use mini-batches (stochastic gradient descent) and optimisers like Adam.

Examples

  • Squared error: Predicted 7, true 10: loss (7 − 10)² = 9.
  • Learning rate: 0.1 might bounce around; 0.0001 might take forever.
  • Mini-batches: Update after every 64 examples instead of the whole dataset.

How it works

  1. Compute the loss on a batch of examples.
  2. Compute the gradient of the loss for every weight.
  3. Step each weight a little in the downhill direction, and repeat.

Check your understanding

Gradient descent moves weights in which direction?
Options: Opposite to the gradient, downhill; Along the gradient, uphill; In a random direction.
Answer: Opposite to the gradient, downhill. The gradient points uphill, so we step the other way.
What happens if the learning rate is far too large?
Options: Training can overshoot and diverge; Training becomes perfect; Nothing changes.
Answer: Training can overshoot and diverge. Huge steps jump past the minimum and can blow up.

Remember

Loss measures error; gradient descent repeatedly steps weights downhill to reduce it.

Talk about it

When have you improved something by making small adjustments and checking?

Go deeper

Adam combines momentum with per-parameter adaptive step sizes. Learning-rate schedules with warm-up and decay, plus weight decay, are standard in large-model training.