Type “dog on a beach” into a modern photo app and it finds your pictures, even though nobody tagged them with that phrase. Text-to-image generators need the same skill in the other direction: they must know which pictures fit a prompt. Both often rely on a model that maps images and sentences into one shared EmbeddingA list of numbers (a vector) that represents a token, word, sentence, or image, learned so that similar things end up with similar vectors.Open in glossary space, where a picture and a good description of it land close together.
The best-known model of this kind is CLIP, introduced by OpenAI (Radford et al., 2021). This lesson trains a miniature version in your browser.
Two encoders, one space
CLIP has two separate networks. An image encoder (the paper tried ResNets and Vision Transformers) turns a picture into a vector. A text encoder (a transformer) turns a caption into a vector of the same length. Neither network knows anything about the other’s input. The only thing connecting them is how they are trained.
Both vectors are scaled to length 1, so the comparison between an image and a caption is their Cosine similarityA measure of how closely two vectors point in the same direction, from -1 (opposite) through 0 (perpendicular) to 1 (same direction), ignoring their lengths.Open in glossary , a number from (opposite directions) to (same direction).
Learning from pairs instead of labels
CLIP was trained on 400 million image and text pairs collected from publicly available sources on the internet, each image paired with text that appeared alongside it. Nobody labeled them with categories. The training signal is simply which caption came with which image.
Take a batch of pairs and compute every image against every caption: an matrix of similarities whose diagonal holds the true pairs. Contrastive learningTraining that pulls the embeddings of matching pairs together and pushes non-matching pairs apart, so the model learns a useful space without category labels.Open in glossary asks each image to pick out its own caption from the whole row, and each caption to pick out its own image from the whole column. As a formula, with temperature :
Each bracket is an ordinary SoftmaxA function that turns a list of real-valued scores into probabilities that are positive and sum to 1, by exponentiating each score and dividing by the total.Open in glossary cross-entropy, once along a row and once along a column. This is often called the InfoNCE loss. Every other item in the batch acts as a negative example, which is one reason CLIP used very large batches of 32,768 pairs.
Below, nine toy pictures (three colors times three shapes) are paired with nine captions. Both encoders start random. Press Train and watch both views.
Training two encoders into one space
Each image and caption becomes a point on a circle. Training pulls matching pairs together and pushes the rest apart.
Image embedding (inner ring)Caption embedding (outer ring), labeled by initials
Try this
- Before training, tap red circle. Its nearest caption is essentially random. Train, then tap it again.
- Watch the matrix while training. The diagonal turns red (high similarity) and the rest cools, because the loss rewards both.
- In the circle, each picture slides toward the outer square with its initials (rc for red circle, and so on). Pairs that share a word, like red circle and red square, still end up neighbors.
- Press Reset, set Temperature to 1, and train again. The pairs still line up, but the loss stalls near 1.4. With similarities capped at 1, a temperature of 1 cannot make the right caption much more likely than the others.
What the miniature leaves out
The demo is honest about the method but small in every other way:
- Its “image features” are five hand-made numbers per picture (how red, green, and blue it is, how round, how many corners). CLIP’s image encoder reads raw pixels.
- Its “text features” count words from a six-word vocabulary. CLIP’s text encoder is a transformer that reads tokens in order.
- Its encoders are single linear maps into two dimensions, so you can see the space. CLIP’s embeddings have hundreds of dimensions.
- Its temperature is a slider. CLIP learns the temperature as a parameter during training, starting from 0.07.
Why one space is so useful
Once images and text live in the same space, many tasks become lookups.
Zero-shot classification. To sort photos into categories, write each category as a caption (“a photo of a dog”, “a photo of a cat”) and pick the caption closest to each image. No category-specific training is needed, which is what Zero-shotDoing a task without any examples of that specific task during training or in the prompt, for example classifying images into categories the model was never trained to label.Open in glossary means here. The CLIP paper reports that this matched the accuracy of the original ResNet-50 on ImageNet without using any of ImageNet’s 1.28 million labeled training images.
Search in both directions. Embed every photo once, embed the query text, and return the nearest photos. Or start from an image and find captions or similar images.
Steering other models. Many image generators use a CLIP-style text encoder to turn the prompt into the vectors that condition the denoising network, and many vision-language models connect a CLIP-style image encoder to a language model.
Key ideas
- Two separate encoders, one for images and one for text, can be trained to share an embedding space using only paired data.
- The contrastive loss makes each true pair more similar than every mismatched pair in the batch, in both directions.
- Temperature scales similarities into logits; too high and the model cannot become confident, so CLIP learns it.
- A shared space turns zero-shot classification and cross-modal search into nearest-neighbor lookups.
- What the space captures is what helps match captions, which leaves gaps in counting and relations.