Contrastive learning, CLIP style

Train an image encoder and a text encoder together until each picture lands next to its caption in one shared space.

AdvancedExplained in One space for images and text

Training two encoders into one space

Each image and caption becomes a point on a circle. Training pulls matching pairs together and pushes the rest apart.

Image embedding (inner ring)Caption embedding (outer ring), labeled by initials

0.30
Training steps0
Loss (symmetric InfoNCE)3.182
Images matched to their caption2 of 9

Try this

  • Press Train and watch the diagonal of the similarity matrix light up.
  • Tap an image before and after training to see which caption it is closest to.
  • Reset, raise the temperature to 1, and train again. Compare where the loss settles.

The features and encoders are deliberately tiny; the loss and the training are the real thing. The lesson explains what the full-size model does differently.

Try "embedding", "softmax", "overfitting", or "backpropagation".