RAG, Step by Step
A language model learns from a huge amount of text, but that text stops at some point in time, and it never included your own documents. Ask a model about your team's vacation policy, and it can only guess. Sometimes it says it does not know. Sometimes it gives a confident answer that is simply wrong.
The fix is surprisingly simple: before asking the question, find the right part of your documents and give it to the model together with the question. This is called retrieval-augmented generation, or RAG. First retrieve, then generate. It is built on embeddings, the lists of numbers from the previous article. Let's go through it one step at a time, with a small team handbook.
Step one: cut the documents into small pieces called chunks. We could hand the whole handbook to the model every time, but real documents are long. A model can only read so much at once, every extra word costs time and money, and a lot of unrelated text can make the answer worse. Change the size of the chunks below.
8 sentences, 4 chunks. The dots show each sentence's topic. Real systems usually split by length, often a few hundred tokens, rather than by sentence count.
There is no perfect size. Chunks that are too small lose their meaning, like a sentence that says "this" without saying what. Chunks that are too big mix topics together. Overlap helps with ideas that get cut at the edge of a chunk, at the cost of storing some text twice.
Step two: find the right chunks. This is called retrieval. Every chunk is turned into an embedding once and stored, often in a special database for vectors. When a question comes in, it becomes an embedding too, and we pick the chunks whose embeddings are the most similar. Pick a question and decide how many chunks to send.
- Vacation
New team members get 20 vacation days a year. Two years later, this goes up to 25 days.
sent1.00 - Remote work
You can work from home up to three days a week. Tell your team lead which days you will be in the office.
0.10 - Expenses
Keep the receipts for travel and meals on work trips. Send them to finance within 30 days.
0.05 - Laptop
Everyone gets a laptop on their first day. The company pays for it, and you return it when you leave.
0.02
Example vectors, chosen by hand. Here each chunk is one section of the handbook.
Notice the question about the dog. The handbook says nothing about dogs, but retrieval still returns the closest chunks. It always does. Retrieval finds the most similar text, not the right answer, so the last step has to handle the case where the answer is missing.
Step three: build the prompt, the full text the model receives. It has a short instruction, the chunks we found and the question. Then the model writes the answer. Turn retrieval off to see what happens when the model has to answer from memory.
Answer the question using only the context below. If the answer is not in the context, say you don't know. Context: [Vacation] New team members get 20 vacation days a year. Two years later, this goes up to 25 days. [Remote work] You can work from home up to three days a week. Tell your team lead which days you will be in the office. Question: How many vacation days do I get?
You get 20 vacation days a year. After two years, it goes up to 25.
Matches the handbook.Example answers, written by hand to show typical behavior. Real models do not always follow the instructions, even when they get the right context.
With the right chunks, the model can answer from the handbook instead of guessing. The instruction to say "I don't know" matters too: it gives the model a way out when the chunks do not contain the answer. Without retrieval, the model has nothing to go on, so it may make up an answer that sounds right.
That is RAG: cut, find, then ask. It is how many "chat with your documents" tools and company assistants work. Keep in mind that it lowers the chance of made-up answers, but it does not remove it. Retrieval can miss the right chunk, and the model can still ignore or misread what it was given. A good answer depends on every step, not only on the model.
If this content helped you, you can buy me a coffee.
You can join the newsletter to be notified of awesome interactive articles and courses about software, design and AI. You will receive at most a few emails per month.