Course 5, lesson 42 of 100, Ages 11+

Cleaning data

Fixing mistakes before training

Like I’m 5

Real data is messy, like a toy box full of broken pieces and odd socks. Cleaning data means fixing mistakes and removing junk before AI learns from it.

The big idea

Raw data often has typos, missing values, duplicates and impossible entries, like a person aged 250 or a temperature of 900 degrees. If the AI learns from these, its predictions get worse.

Data scientists often spend more time cleaning data than training models. They remove duplicates, fix formats, fill or flag missing values, and check outliers to see if they're errors or genuinely unusual cases.

Examples

  • Typos: 'Mumbai', 'mumbai' and 'Mumbia' should become one city.
  • Missing values: A blank age can be flagged instead of guessed wildly.
  • Duplicates: The same survey answered twice would count double.

How it works

  1. Look for missing, duplicate or impossible values.
  2. Fix or remove the problems, and note what you changed.
  3. Check the cleaned data still represents reality.

Check your understanding

Which value most likely needs cleaning?
Options: A person's age recorded as 250; An age of 12; An age of 45.
Answer: A person's age recorded as 250. Nobody is 250 years old, so it's probably an error.
Why should you note what you changed while cleaning?
Options: So others can check and repeat your work; To make the file bigger; It's not needed.
Answer: So others can check and repeat your work. Recording changes keeps the work honest and repeatable.

Remember

Clean data before training: fix errors, remove duplicates and record what you changed.

Talk about it

What mistakes might appear in a class attendance list?

Go deeper

Common steps include deduplication, normalising formats, imputing missing values and outlier detection. Cleaning choices can introduce bias, so they should be documented.