Course 10, lesson 95 of 100, Adults

Scaling laws

Why bigger models got better

Like I’m 5

Scientists found that if you make an AI bigger, give it more to read and more computer time, it gets better in a surprisingly predictable way.

The big idea

Scaling laws describe how a model's loss falls as a smooth power law when you increase parameters, training data and compute. This let labs predict the performance of a big model from small experiments before spending millions.

Later work (the 'Chinchilla' study, 2022) showed many models were undertrained: for a fixed compute budget, it's better to balance size and data, roughly 20 training tokens per parameter. Bigger models also show new abilities, though whether these appear suddenly or gradually is debated.

Examples

  • Prediction: Small runs forecast how a 10× bigger model will score.
  • Balance: A smaller model trained on more data can beat a bigger, undertrained one.
  • Compute: Training compute is measured in floating-point operations (FLOPs).

How it works

  1. Train small models across a range of sizes and data amounts.
  2. Fit a power-law curve to their losses.
  3. Use it to choose the best size and data for a big training run.

Check your understanding

What do scaling laws describe?
Options: How loss falls predictably as models, data and compute grow; How fast computers boot; The size of screens.
Answer: How loss falls predictably as models, data and compute grow. They're empirical power laws linking scale to performance.
What did the Chinchilla study suggest?
Options: Balance model size and training data for a fixed budget; Always build the biggest model possible; Data doesn't matter.
Answer: Balance model size and training data for a fixed budget. Roughly 20 tokens per parameter was compute-optimal in that study.

Remember

Loss falls predictably with scale, and data should grow along with model size.

Talk about it

What else in life gets predictably better with more practice and resources?

Go deeper

Kaplan et al. (2020) and Hoffmann et al. (2022) are the key papers. Inference-time compute, such as longer reasoning, has emerged as another axis of scaling.