Ask a language model about your company’s return policy and it can only guess. It never saw the policy: its knowledge stopped when training ended, and your documents were never part of that training. Worse, it may guess fluently, because a language model always produces plausible text whether or not it has the facts. Confident, invented answers like this are called HallucinationWhen a language model states something false or invented, such as a fake quote or citation, in the same fluent, confident style as true statements.Open in glossary.
Retrieval-augmented generationA technique that searches a collection of documents for passages relevant to a question and adds them to a language model's prompt, so it can answer from that text.Open in glossary (RAG) is the most common way to reduce this. Before the model answers, a search step finds the passages most relevant to the question, and they are pasted into the prompt. The model then answers from text it can actually see. The name comes from a 2020 paper by Lewis and colleagues at Facebook AI Research; the pattern now underlies most assistants that answer questions about documents.
The pipeline
Everything you need comes from the last two lessons:
- Chunk. Split each document into passages small enough to embed and to fit in a prompt.
- Embed and store. Turn each chunk into a vector with a sentence embedding model and store it, usually in a vector database.
- Retrieve. When a question arrives, embed it and find the k chunks with the highest cosine similarity.
- Generate. Put those chunks and the question into a prompt, with instructions to answer from the given text, and send it to a language model.
Steps 1 and 2 happen once, ahead of time. Steps 3 and 4 happen for every question.
Retrieval-augmented generation, step by step
Answer questions about a document the model has never seen: split it, embed it, find the relevant pieces, and put them in the prompt.
1 Split the document into chunks
Whole sentences are packed together up to the chunk size.
- #1Northwind Cycles staff handbook. This shop and handbook are fictional and exist only for this lesson. Opening hours. The shop opens at 9 am and closes at 6 pm from Tuesday to Saturday. It is closed on Sundays and Mondays.
- #2On public holidays the shop opens from 10 am to 3 pm. Returns. Customers may return an unused bike within 30 days of purchase for a full refund. Used bikes can be exchanged within 14 days but not refunded.
- #3Clothing and helmets must be returned in their original packaging. Warranty. Every new frame carries a five-year warranty against manufacturing defects. Wheels, brakes, and other parts are covered for one year.
- #4The warranty does not cover damage from crashes, racing, or ordinary wear. Repairs. A basic tune-up costs 45 dollars and takes one working day. Customers can book a repair online or at the counter.
- #5Bikes left more than 60 days after a repair is finished may be donated to the community bike program. Staff benefits.
- #6Employees receive a 25 percent discount on bikes and a 40 percent discount on parts and clothing after their first three months. Staff may borrow a demo bike for up to one week, twice a year. Bike fitting.
- #7A professional fitting takes about 90 minutes and costs 120 dollars. Customers who buy a bike within two weeks of a fitting get the fitting fee back as store credit.
Runs a real model in your browser
Embed the chunks and your question to find the most relevant passages. This uses all-MiniLM-L6-v2, a small sentence-embedding model that maps text to 384 numbers.
The first time, your browser downloads about 24 MB of model files from Hugging Face, plus a 14 MB runtime from this site. Both are cached for later visits. Everything runs on your device; nothing you type is sent anywhere.
Try this
- With the default question about bringing back a bike, check that the top chunk mentions returns. The question says “bring back”; the handbook says “return”. The embedding model connects them anyway.
- Try Do employees get money off? The handbook says “discount”, never “money off”.
- Drag Chunk size down to 10. Chunks shrink to single sentences, and some, such as “Wheels, brakes, and other parts are covered for one year”, no longer say what they belong to. Longer sentences are cut in two: ask the warranty question and the best match ends mid-sentence. Drag it up to 120 and each chunk mixes several topics.
- Set Chunks to retrieve to 1 and ask If I crash my bike, will the warranty pay?. Then raise it to 3. More chunks make it likelier the answer is in the prompt, but also add text that has nothing to do with the question.
- Read the assembled prompt. Everything the model will know about this shop is in that box.
This lab stops at the prompt, because the generation step is the part you already know: a large language model continues the text. The interesting decisions are all upstream, in what gets retrieved.
Choosing chunk size and k
ChunkingSplitting documents into smaller passages before embedding them, so each passage can be retrieved and fit in a prompt on its own.Open in glossary is a tradeoff. Small chunks give precise matches but can lose the context that makes them meaningful. Large chunks keep context but blend topics, so their single vector matches everything a little and nothing well. Many systems split on natural boundaries such as sentences, paragraphs, or headings, keep some overlap between neighboring chunks, and attach the document title to each chunk.
The number of chunks retrieved, k, trades recall against noise and cost. Every retrieved chunk spends part of the model’s Context windowThe maximum number of tokens a language model can take into account at once, covering both the prompt and the text it generates.Open in glossary, and irrelevant passages can distract the model. Production systems often retrieve generously and then use a second, more accurate model, a reranker, to keep only the best few.
What RAG does not fix
- Retrieval can miss. If the relevant passage is not in the top k, the model never sees it and may answer anyway. Good systems tell the model to say “I don’t know” when the context does not contain the answer, as the prompt in the demo does.
- The model can ignore or misread the context. Grounding makes errors less likely, not impossible. Showing sources lets people check.
- Retrieved text is untrusted input. A document can contain instructions, such as “ignore your previous instructions”, and the model may follow them. This is called Prompt injectionAn attack in which text the model reads, such as a web page, email, or document, contains instructions that override or redirect what the user or developer intended.Open in glossary, and any system that retrieves web pages, emails, or user uploads has to defend against it.
- The index has to stay current. Answers are only as fresh and as accurate as the documents that were chunked and embedded.
Key ideas
- RAG supplies a language model with relevant passages at question time, so it can use private or recent information it was never trained on.
- The pipeline is chunk, embed and store, retrieve the top k, then generate from a prompt that includes the retrieved text.
- Chunk size and k are tradeoffs between precision, context, noise, and cost.
- Retrieval improves grounding but does not guarantee correct answers, and retrieved text must be treated as untrusted input.