Training two encoders into one space
Each image and caption becomes a point on a circle. Training pulls matching pairs together and pushes the rest apart.
Image embedding (inner ring)Caption embedding (outer ring), labeled by initials
Training steps0
Loss (symmetric InfoNCE)3.182
Images matched to their caption2 of 9
Try this
- Press Train and watch the diagonal of the similarity matrix light up.
- Tap an image before and after training to see which caption it is closest to.
- Reset, raise the temperature to 1, and train again. Compare where the loss settles.
The features and encoders are deliberately tiny; the loss and the training are the real thing. The lesson explains what the full-size model does differently.