Many of today’s image generators, including Stable Diffusion, are Diffusion modelA generative model that learns to reverse a gradual noising process, so it can turn pure random noise into new data such as images, step by step.Open in glossary. Given a prompt, they do not paint from left to right. They start from a canvas of pure random noise and, over a few dozen steps, remove a little noise at a time until a picture is left. This lesson shows why that works, using data simple enough that the hard part can be computed exactly.
Destroying data is easy
Start with the easy direction. Take a clean data point and blend it with Gaussian noise :
Here runs from 0 to 1 and (read “alpha bar”) falls smoothly from 1 to nearly 0. At you have the data. At you have noise that no longer remembers where it came from. Because the signal shrinks while the noise grows, a sample with unit variance keeps unit variance throughout, which is why this is called a variance-preserving process. This is the forward process of DDPM (Ho et al., 2020), written in the continuous-time form of Song et al. (2021); the demo uses their schedule, with rising linearly from 0.1 to 20 and .
Below, every dot is a sample from a 2D dataset made of small Gaussian blobs arranged in a shape. Set the direction to Add noise first.
Diffusion on a 2D dataset
Dots are samples. Blue shading is the noisy data density at time t; arrows show its score, the direction toward higher density.
Try this
- Choose Add noise and drag Progress slowly. The shape blurs, shrinks toward the center, and by the end is a round cloud of noise. Every dataset ends in the same cloud.
- Switch to Remove noise and press Denoise. The same kind of cloud reorganizes into the shape. Little happens for most of the run; the shape snaps into focus near the end, when the noise level is small.
- Watch the arrows while scrubbing. Early on they all point loosely toward the middle. Late in the run they point sharply toward the nearest blob.
- Compare Deterministic (ODE) with Random (SDE): two ways to run the process backward, one following smooth paths and one adding a little fresh noise at each step (both explained below). Both end on the data, but the random sampler’s paths jitter and individual dots end in different places.
Running time backward
Destroying structure takes no knowledge at all. Rebuilding it seems to require knowing what the data looks like. The surprising result behind diffusion models is that you need exactly one thing: the ScoreThe gradient of the log of a probability density, which points in the direction where data becomes more likely. Diffusion models learn the score of noisy data at every noise level.Open in glossary of the noisy data distribution at every time,
the direction in which noisy data becomes more likely. That is what the arrows draw. With the score in hand, there are two standard ways to run the process in reverse, from down to :
is the reverse-time stochastic differential equation (a result due to Anderson, 1982, used for generation by Song et al., 2021). It adds fresh randomness at each step, which is the Random (SDE) option. Removing the random term and halving the score term gives the probability flow ODE,
whose paths are deterministic but whose samples, taken together, have exactly the same distribution at every time. That is the Deterministic (ODE) option. In both, the term undoes the forward shrinking and the score term pulls samples toward likely regions.
What this demo knows that a real model does not
This demo cheats, openly. Its data is a mixture of Gaussians, and adding Gaussian noise to a mixture of Gaussians gives another mixture of Gaussians with wider, shrunken blobs. So and its score have exact formulas, and the arrows you see are the true score.
For images nobody knows . A diffusion model therefore trains a neural network to approximate the score. In DDPM the network is trained to predict the noise that was mixed in, with a simple loss:
Predicting the noise and estimating the score are the same task in disguise: the best possible noise prediction satisfies . Training needs only clean examples and the forward formula, which is why diffusion models are comparatively easy to train.
The exact score of a noisy Gaussian mixtureOptional
If the data is , then
Write and for those means and variances. The score is a weighted average of each blob’s own score:
where is the probability that came from blob . The site’s tests check this formula against finite differences of , and check that both samplers turn standard normal noise back into the data distribution.
From dots to pictures
Nothing above depends on the data being two-dimensional. A 512 by 512 color image is 786,432 numbers, and the forward process, the score, and both samplers work the same way in that space. Practical systems add a few ideas on top:
- Latent diffusion (Rombach et al., 2022), used by Stable Diffusion, first compresses images with an autoencoder and runs diffusion in the smaller compressed Latent spaceThe space of learned, hidden coordinates a model uses to represent things internally, such as word vectors or a network's hidden layers. Nearby points stand for similar things.Open in glossary, which is much cheaper.
- Conditioning feeds the prompt’s text embedding into the noise-prediction network, so the network predicts the noise for “a photo of a cat” rather than for images in general.
- Classifier-free guidanceA sampling technique for conditional diffusion models that compares the model's prediction with and without the prompt and moves further in the prompt's direction, trading variety for closer prompt matching.Open in glossary (Ho and Salimans, 2022) runs the network with and without the prompt and pushes the result further in the prompt’s direction, trading variety for closer prompt matching.
- Fewer steps. The plain reverse process can take hundreds of steps. Better ODE solvers and distillation methods cut this to a few dozen or even a handful.
Key ideas
- The forward process mixes data with Gaussian noise, , until only noise remains.
- If you know the score at every noise level, you can run the process backward, with a random (SDE) or deterministic (ODE) sampler.
- Real models do not know the score. They train a network to predict the added noise, which is equivalent to estimating the score.
- Structure is rebuilt coarse to fine: overall shape first, details last.
- Image generators apply the same mathematics in many more dimensions, usually in a compressed latent space and steered by the prompt.