Course 10, lesson 98 of 100, Adults
How diffusion models work
From noise to images
Like I’m 5
Imagine slowly adding fog to a picture until it's all fog. A diffusion model learns to do the opposite: remove the fog a little at a time until a picture appears.
The big idea
Training uses a forward process that adds a little Gaussian noise to images over many steps until they become pure noise. A neural network learns to predict the noise added at each step, given the noisy image, the step number and the text prompt.
To generate, start from random noise and repeatedly subtract the predicted noise. Latent diffusion does this in a compressed space for speed. Classifier-free guidance strengthens how closely the image follows the prompt, at some cost to variety.
Examples
- Training: Learn to undo noise at every level, from a light haze to pure static.
- Latent space: Stable Diffusion works on a small compressed version of the image.
- Guidance: Higher guidance follows the prompt more strictly.
How it works
- Start from random noise.
- Predict and remove a little noise, guided by the prompt.
- Repeat many times until a clear image remains.
Check your understanding
- What does the network learn to predict during training?
- Options: The noise added to an image; The image's file name; The camera brand.
Answer: The noise added to an image. Predicting noise lets it reverse the noising process. - Why do latent diffusion models work in a compressed space?
- Options: It's much faster and cheaper; It makes images blurry on purpose; It's required by law.
Answer: It's much faster and cheaper. Smaller representations cut compute dramatically.
Remember
Diffusion models learn to remove noise step by step, turning static into images.
Talk about it
What other processes could you run backwards, like un-mixing paint?
Go deeper
Key papers include DDPM (Ho et al., 2020) and latent diffusion (Rombach et al., 2022). Text conditioning typically uses cross-attention to text embeddings; faster samplers reduce the number of steps.