Lesson 1 of 5

Diffusion: from noise to data

Image generators start from pure noise and remove it a little at a time. Watch the exact reverse process turn random dots into a dataset, and see what a real model has to learn.

Advanced22 min

In this lesson you will

  • Describe the forward process that gradually turns data into noise
  • Explain what the score is and why knowing it lets you run the noising backward
  • Compare a deterministic and a random sampler on the same data
  • Connect the toy demo to how real diffusion models are trained

Many of today’s image generators, including Stable Diffusion, are Diffusion modelA generative model that learns to reverse a gradual noising process, so it can turn pure random noise into new data such as images, step by step.Open in glossary. Given a prompt, they do not paint from left to right. They start from a canvas of pure random noise and, over a few dozen steps, remove a little noise at a time until a picture is left. This lesson shows why that works, using data simple enough that the hard part can be computed exactly.

Destroying data is easy

Start with the easy direction. Take a clean data point x0x_0 and blend it with Gaussian noise ε∼N(0,I)\varepsilon \sim \mathcal{N}(0, I):

xt=αˉt x0+1−αˉt εx_t = \sqrt{\bar\alpha_t}\, x_0 + \sqrt{1 - \bar\alpha_t}\, \varepsilon

Here tt runs from 0 to 1 and αˉt\bar\alpha_t (read “alpha bar”) falls smoothly from 1 to nearly 0. At t=0t = 0 you have the data. At t=1t = 1 you have noise that no longer remembers where it came from. Because the signal shrinks while the noise grows, a sample with unit variance keeps unit variance throughout, which is why this is called a variance-preserving process. This is the forward process of DDPM (Ho et al., 2020), written in the continuous-time form of Song et al. (2021); the demo uses their schedule, with β(t)\beta(t) rising linearly from 0.1 to 20 and αˉt=e−∫0tβ(s) ds\bar\alpha_t = e^{-\int_0^t \beta(s)\,ds}.

Below, every dot is a sample from a 2D dataset made of small Gaussian blobs arranged in a shape. Set the direction to Add noise first.

Diffusion on a 2D dataset

Dots are samples. Blue shading is the noisy data density at time t; arrows show its score, the direction toward higher density.

Direction
Dataset
Sampler
t = 1.000
Time t1.000
Signal kept, sqrt(abar)0.007
Noise, sqrt(1 - abar)1.000

Try this

  • Choose Add noise and drag Progress slowly. The shape blurs, shrinks toward the center, and by the end is a round cloud of noise. Every dataset ends in the same cloud.
  • Switch to Remove noise and press Denoise. The same kind of cloud reorganizes into the shape. Little happens for most of the run; the shape snaps into focus near the end, when the noise level is small.
  • Watch the arrows while scrubbing. Early on they all point loosely toward the middle. Late in the run they point sharply toward the nearest blob.
  • Compare Deterministic (ODE) with Random (SDE): two ways to run the process backward, one following smooth paths and one adding a little fresh noise at each step (both explained below). Both end on the data, but the random sampler’s paths jitter and individual dots end in different places.

Running time backward

Destroying structure takes no knowledge at all. Rebuilding it seems to require knowing what the data looks like. The surprising result behind diffusion models is that you need exactly one thing: the ScoreThe gradient of the log of a probability density, which points in the direction where data becomes more likely. Diffusion models learn the score of noisy data at every noise level.Open in glossary of the noisy data distribution at every time,

s(x,t)=∇xlog⁡pt(x),s(x, t) = \nabla_x \log p_t(x),

the direction in which noisy data becomes more likely. That is what the arrows draw. With the score in hand, there are two standard ways to run the process in reverse, from t=1t = 1 down to t=0t = 0:

dx=[−12β(t) x−β(t) s(x,t)] dt+β(t) dWˉdx = \Big[-\tfrac{1}{2}\beta(t)\, x - \beta(t)\, s(x, t)\Big]\,dt + \sqrt{\beta(t)}\, d\bar{W}

is the reverse-time stochastic differential equation (a result due to Anderson, 1982, used for generation by Song et al., 2021). It adds fresh randomness at each step, which is the Random (SDE) option. Removing the random term and halving the score term gives the probability flow ODE,

dxdt=−12β(t) [x+s(x,t)],\frac{dx}{dt} = -\tfrac{1}{2}\beta(t)\,\big[x + s(x, t)\big],

whose paths are deterministic but whose samples, taken together, have exactly the same distribution at every time. That is the Deterministic (ODE) option. In both, the xx term undoes the forward shrinking and the score term pulls samples toward likely regions.

What this demo knows that a real model does not

This demo cheats, openly. Its data is a mixture of Gaussians, and adding Gaussian noise to a mixture of Gaussians gives another mixture of Gaussians with wider, shrunken blobs. So ptp_t and its score have exact formulas, and the arrows you see are the true score.

For images nobody knows ptp_t. A diffusion model therefore trains a neural network to approximate the score. In DDPM the network εθ(xt,t)\varepsilon_\theta(x_t, t) is trained to predict the noise that was mixed in, with a simple loss:

L=Ex0, ε, t[∥ε−εθ(xt,t)∥2]\mathcal{L} = \mathbb{E}_{x_0,\, \varepsilon,\, t}\Big[\big\lVert \varepsilon - \varepsilon_\theta(x_t, t) \big\rVert^2\Big]

Predicting the noise and estimating the score are the same task in disguise: the best possible noise prediction satisfies s(xt,t)=− E[ε∣xt]/1−αˉts(x_t, t) = -\,\mathbb{E}[\varepsilon \mid x_t] / \sqrt{1 - \bar\alpha_t}. Training needs only clean examples and the forward formula, which is why diffusion models are comparatively easy to train.

The exact score of a noisy Gaussian mixtureOptional

If the data is ∑kπk N(μk,σk2I)\sum_k \pi_k\, \mathcal{N}(\mu_k, \sigma_k^2 I), then

pt(x)=∑kπk N(x; αˉt μk, (αˉtσk2+1−αˉt) I).p_t(x) = \sum_k \pi_k\, \mathcal{N}\big(x;\ \sqrt{\bar\alpha_t}\,\mu_k,\ (\bar\alpha_t \sigma_k^2 + 1 - \bar\alpha_t)\, I\big).

Write mkm_k and vkv_k for those means and variances. The score is a weighted average of each blob’s own score:

∇xlog⁡pt(x)=∑krk(x) mk−xvk,rk(x)=πk N(x;mk,vkI)pt(x),\nabla_x \log p_t(x) = \sum_k r_k(x)\, \frac{m_k - x}{v_k}, \qquad r_k(x) = \frac{\pi_k\, \mathcal{N}(x; m_k, v_k I)}{p_t(x)},

where rk(x)r_k(x) is the probability that xx came from blob kk. The site’s tests check this formula against finite differences of log⁡pt\log p_t, and check that both samplers turn standard normal noise back into the data distribution.

From dots to pictures

Nothing above depends on the data being two-dimensional. A 512 by 512 color image is 786,432 numbers, and the forward process, the score, and both samplers work the same way in that space. Practical systems add a few ideas on top:

  • Latent diffusion (Rombach et al., 2022), used by Stable Diffusion, first compresses images with an autoencoder and runs diffusion in the smaller compressed Latent spaceThe space of learned, hidden coordinates a model uses to represent things internally, such as word vectors or a network's hidden layers. Nearby points stand for similar things.Open in glossary, which is much cheaper.
  • Conditioning feeds the prompt’s text embedding into the noise-prediction network, so the network predicts the noise for “a photo of a cat” rather than for images in general.
  • Classifier-free guidanceA sampling technique for conditional diffusion models that compares the model's prediction with and without the prompt and moves further in the prompt's direction, trading variety for closer prompt matching.Open in glossary (Ho and Salimans, 2022) runs the network with and without the prompt and pushes the result further in the prompt’s direction, trading variety for closer prompt matching.
  • Fewer steps. The plain reverse process can take hundreds of steps. Better ODE solvers and distillation methods cut this to a few dozen or even a handful.

Key ideas

  • The forward process mixes data with Gaussian noise, xt=αˉt x0+1−αˉt εx_t = \sqrt{\bar\alpha_t}\,x_0 + \sqrt{1-\bar\alpha_t}\,\varepsilon, until only noise remains.
  • If you know the score ∇xlog⁡pt(x)\nabla_x \log p_t(x) at every noise level, you can run the process backward, with a random (SDE) or deterministic (ODE) sampler.
  • Real models do not know the score. They train a network to predict the added noise, which is equivalent to estimating the score.
  • Structure is rebuilt coarse to fine: overall shape first, details last.
  • Image generators apply the same mathematics in many more dimensions, usually in a compressed latent space and steered by the prompt.

Check yourself

Pick an answer to see why it is right or wrong. Nothing is graded. Your first answer is saved in this browser so the question can come back for review.

1In the forward process at time t, what is a noisy sample made of?
2What does the score of the noisy distribution tell the sampler?
3Why can this demo compute the score exactly when an image model cannot?

Progress is saved in this browser only.

Up nextOne space for images and text
Next
Frontiers
  1. 1Diffusion: from noise to data
  2. 2One space for images and text
  3. 3Mixture of experts
  4. 4Thinking longer
  5. 5Models that use tools

Try "embedding", "softmax", "overfitting", or "backpropagation".