Course 10, lesson 96 of 100, Adults

RLHF and preference tuning

Teaching models what people prefer

Like I’m 5

After an AI learns to write, people show it which answers they like better. Bit by bit, it learns to give more helpful and kinder answers.

The big idea

Reinforcement learning from human feedback starts with a supervised model. People compare pairs of responses and pick the better one; a reward model learns to predict those preferences. The language model is then optimised to score highly, with a penalty for drifting too far from its original behaviour.

Simpler alternatives like Direct Preference Optimisation train on preference pairs directly. Constitutional AI uses a written set of principles and AI-generated feedback to scale the process. All of these improve helpfulness and safety, but can also teach models to tell people what they want to hear, so evaluation matters.

Examples

  • Comparisons: Raters choose between two answers to the same prompt.
  • Reward model: A model that predicts which answer people would prefer.
  • Sycophancy: Over-optimising for approval can make a model too agreeable.

How it works

  1. Collect human comparisons between pairs of answers.
  2. Train a reward model to predict the preferred answer.
  3. Optimise the language model towards higher reward, staying close to the original.

Check your understanding

What does the reward model learn in RLHF?
Options: To predict which answers people prefer; To generate images; To count tokens.
Answer: To predict which answers people prefer. It turns human comparisons into a learnable score.
What's a known risk of optimising for human approval?
Options: Sycophancy: telling people what they want to hear; Models become too slow; Models forget the alphabet.
Answer: Sycophancy: telling people what they want to hear. Approval isn't always the same as truth, so evals check for it.

Remember

Preference tuning uses human or AI comparisons to steer models towards helpful, safe answers.

Talk about it

If you rated chatbot answers, what would make one answer better than another?

Go deeper

InstructGPT (Ouyang et al., 2022) popularised RLHF with PPO and a KL penalty. DPO (Rafailov et al., 2023) and Constitutional AI (Bai et al., 2022) are widely used alternatives.