Course 2, lesson 19 of 100, Ages 6+
Learning from rewards
Points for good moves
Like I’m 5
When you train a puppy, you give it a treat when it sits. The puppy learns that sitting is good. Some AI learns the same way, with points instead of treats.
The big idea
In reward learning, an AI tries actions and gets a score: points for good results and lost points for bad ones. Nobody tells it the right move. It has to discover it.
Over many tries, it learns which actions earn more reward. That's how AI learned to balance robots and play video games. Picking the right reward matters: reward the wrong thing and you get the wrong behaviour.
Examples
- Puppy training: Treats for sitting means more sitting.
- Video game AI: Points for finishing a level faster.
- A silly reward: A boat-racing AI learned to spin in circles collecting points instead of finishing the race.
How it works
- The AI tries an action.
- It gets a reward or a penalty.
- It does more of whatever earned rewards.
Check your understanding
- How does reward learning teach an AI?
- Options: It earns points for good results and learns to get more; It reads a rule book; It copies the answers.
Answer: It earns points for good results and learns to get more. Rewards guide it towards better actions over many tries. - Why must people choose rewards carefully?
- Options: The AI will chase whatever earns points, even silly tricks; Rewards cost money; AI dislikes rewards.
Answer: The AI will chase whatever earns points, even silly tricks. AI optimises exactly what it's rewarded for, so the reward must match what we really want.
Remember
Reward learning means try, get points, and do more of what works. Choose the reward wisely.
Talk about it
What reward would you give a robot that tidies your room?
Go deeper
This is reinforcement learning. An agent takes actions in an environment and learns a policy that maximises total reward. When it exploits a badly chosen reward, that's called reward hacking.