Course 6, lesson 53 of 100, Ages 12+

How giant models learn

Read, practise, get feedback

Like I’m 5

First the AI reads a huge library. Then it practises with a teacher. Then people give it gold stars for good answers.

The big idea

Big language models are trained in stages. First, pre-training: reading a huge library of text to learn language and facts by predicting the next token.

Then fine-tuning teaches the model to follow instructions and chat helpfully, using example conversations written by people. Finally, feedback training rewards answers people rate as helpful, honest and harmless.

Examples

  • Pre-training: Reading books, websites and code for months on huge computers.
  • Instruction tuning: Learning from examples like 'Q: summarise this. A: ...'.
  • Feedback: People compare two answers and choose the better one.

How it works

  1. Big AI models first read an enormous amount of text. This is called pre-training.
  2. Then they practise being helpful on examples written by people. That’s fine-tuning.
  3. Finally, people rate their answers, and the model learns which answers are better.

Check your understanding

What happens in pre-training?
Options: The model reads lots of text; The model goes to school; The model gets a new name.
Answer: The model reads lots of text. Pre-training is when the model learns the patterns of language from huge amounts of text.
What happens during pre-training?
Options: The model reads huge amounts of text and predicts next tokens; People chat with it every day; It learns to drive.
Answer: The model reads huge amounts of text and predicts next tokens. Pre-training builds general language skill from massive text.

Remember

Big models learn by reading, practising and getting feedback.

Talk about it

What’s the biggest book you’ve ever read?

Go deeper

Large language models are pre-trained to predict the next token across vast collections of text, then refined with supervised fine-tuning and feedback-based methods such as reinforcement learning from human feedback (RLHF). Training the biggest models takes enormous computing power and energy.