Lesson 5 of 6

Search by meaning

A sentence embedding model turns a whole sentence into one vector, so you can search by what text means instead of which words it uses. Run a real one in your browser and race it against keyword search.

Intermediate15 min

In this lesson you will

  • Explain what a sentence embedding is and how it is used for search
  • Compare keyword search with embedding search and predict where each wins
  • Describe how search scales to millions of documents

Word embeddings give each word one vector. Search needs more: one vector for a whole sentence, a question, or a page, so that “how can I save cash” lands near “putting a little money aside every month” even though they share no words. That is what a Sentence embeddingA single vector representing a whole sentence or passage, produced by a model trained so that texts with similar meaning get similar vectors.Open in glossary model does.

One vector per sentence

A sentence embedding model is a transformer that reads the whole text, produces a vector for every token, and then pools them into one vector, often by averaging (some models use the first or last token’s vector instead). The model is trained so that texts with similar meaning get similar vectors, typically by showing it pairs that belong together (a question and its answer, a sentence and its paraphrase) and pulling those vectors together while pushing unrelated pairs apart. This kind of training is called Contrastive learningTraining that pulls the embeddings of matching pairs together and pushes non-matching pairs apart, so the model learns a useful space without category labels.Open in glossary.

The demo below runs all-MiniLM-L6-v2, a widely used open model from the sentence-transformers project. It is small, six transformer layers, and was fine-tuned with a contrastive objective on more than a billion sentence pairs. It turns any text into 384 numbers, scaled to length 1, so cosine similarity is simply the dot product. It was trained on texts up to 256 tokens long, and longer texts are cut off (this browser version cuts at 512), which is one reason long documents are split into chunks before embedding.

Semantic searchSearch that ranks documents by how similar their meaning is to the query, usually by comparing sentence embeddings with cosine similarity.Open in glossary embeds every document once, embeds each query when it arrives, and ranks documents by cosine similarity. Keyword search, the approach behind classic search engines, ranks documents by the words they share with the query. The standard keyword method is BM25A standard keyword ranking formula. It scores documents by the query words they contain, weighting rare words more and limiting the reward for repeated words.Open in glossary, which rewards documents containing the query’s words, gives more weight to rare words than to common ones, and gives less and less extra credit each time a word repeats.

Keyword search versus search by meaning

The same question, two ways. Keyword search matches words; meaning search compares embeddings.

Example searches

Keyword search (BM25)

Scores sentences that contain your exact words, weighting rare words more.

No sentence contains any of your words, so keyword search finds nothing.

Search by meaning (embeddings)

Embeds every sentence and your query, then ranks by cosine similarity.

Runs a real model in your browser

Embed each sentence and your query to rank by meaning. This uses all-MiniLM-L6-v2, a small sentence-embedding model that maps text to 384 numbers.

The first time, your browser downloads about 24 MB of model files from Hugging Face, plus a 14 MB runtime from this site. Both are cached for later visits. Everything runs on your device; nothing you type is sent anywhere.

Try this

  • Before downloading the model, run the examples through keyword search alone. “how can I save cash” finds nothing: no sentence contains “save” or “cash”.
  • Download the model and run the same query. The sentence about putting money aside comes first by a wide margin, and the one about cooking at home being cheaper comes second.
  • Try “animals that like to climb”. Keyword search misses “climbing” because it matches exact words only; meaning search finds the cats.
  • Try “what time can I visit the gallery”. Keyword search grabs the one sentence containing “time”, which is about saving money. Meaning search finds the museum.
  • Try “stock market”. When the exact words matter, both methods agree.
  • Turn on Edit the sentences, add your own, and try to write a query that fools the embedding model.

Neither method is simply better. Embedding search handles paraphrases and synonyms, but it can miss exact strings such as product codes, names, and rare technical terms, and it will always return something, even for a query nothing in the collection answers. Keyword search is precise and easy to explain but blind to meaning. Production search systems often run both and combine the scores, an approach called hybrid search.

Searching millions of vectors

The demo compares the query with every sentence, which takes no time for twelve sentences. For millions of documents, systems store embeddings in a Vector databaseA database built to store embeddings and quickly find the ones most similar to a query vector, usually with an approximate nearest neighbor index.Open in glossary with an approximate nearest neighbor index. A popular one, HNSW (hierarchical navigable small world graphs, Malkov and Yashunin), links each vector to a few of its neighbors in a layered graph and finds a query’s neighbors by walking the graph, checking only a tiny fraction of the stored vectors. The result is almost always the same as an exhaustive search, at a fraction of the cost.

Key ideas

  • A sentence embedding model maps a whole text to one vector, trained so that similar meanings get similar vectors.
  • Semantic search ranks documents by cosine similarity to the query’s embedding.
  • Keyword search (BM25) matches exact words and wins on names and codes; embedding search wins on paraphrases. Many systems combine both.
  • Approximate nearest neighbor indexes make vector search fast at the scale of millions or billions.

Check yourself

Pick an answer to see why it is right or wrong. Nothing is graded. Your first answer is saved in this browser so the question can come back for review.

1You search a help center for "cancel my plan" and the right article is titled "Ending your subscription". Which search is more likely to find it?
2Someone searches a parts catalog for the exact code "XR-2270B". Which statement is most accurate?
3How do search systems find the most similar vectors among millions without comparing against every one?

Progress is saved in this browser only.

Up nextRetrieval-augmented generation
Next
Language as vectors
  1. 1Tokens
  2. 2Vectors and similarity
  3. 3Word embeddings
  4. 4Seeing high dimensions
  5. 5Search by meaning
  6. 6Retrieval-augmented generation

Try "embedding", "softmax", "overfitting", or "backpropagation".