Course 10, lesson 99 of 100, Adults
Interpretability
Looking inside the black box
Like I’m 5
Scientists are building microscopes for AI, so they can see which ideas light up inside a model when it reads or writes.
The big idea
Interpretability research tries to understand what models represent and how they compute. Probes test whether information, like sentiment or position, can be read from internal activations. Attribution methods show which inputs mattered.
Mechanistic interpretability goes further, finding 'features' (directions in activation space that correspond to concepts) and 'circuits' that connect them. Models pack more features than neurons (superposition), so tools like sparse autoencoders are used to pull them apart. This matters for safety: you can't fully trust what you can't inspect.
Examples
- Probes: A simple classifier reads 'is this sentence positive?' from hidden states.
- Features: A direction that activates for a specific concept, like a famous landmark.
- Steering: Turning a feature up or down changes the model's behaviour.
How it works
- Record the model's internal activations on many inputs.
- Find features or circuits linked to specific concepts or behaviours.
- Test them by intervening, turning features up or down.
Check your understanding
- What is superposition in neural networks?
- Options: Packing more features than there are neurons; Two models running together; A type of GPU.
Answer: Packing more features than there are neurons. Features share neurons, so special tools are needed to separate them. - Why does interpretability matter for safety?
- Options: It helps check what a model is really doing; It makes models faster; It removes the need for testing.
Answer: It helps check what a model is really doing. Understanding internals helps detect hidden problems.
Remember
Interpretability looks inside models to find the features and circuits behind their behaviour.
Talk about it
If you could look inside an AI's 'mind', what would you check first?
Go deeper
Influential work includes circuits research on vision models, 'Toy Models of Superposition' (Elhage et al., 2022) and dictionary learning with sparse autoencoders on language models (2023 to 2024).