Course 9, lesson 83 of 100, Ages 14+
RAG: answers from your documents
Retrieval-augmented generation
Like I’m 5
Before answering, the AI looks things up in a trusted folder of notes, then answers using what it found. It's like checking your textbook before answering a question.
The big idea
Retrieval-augmented generation (RAG) combines search with an LLM. Your documents are split into chunks and embedded. When someone asks a question, the closest chunks are retrieved and placed in the prompt, and the model answers using them.
RAG keeps answers grounded in up-to-date, private information without retraining the model, and it lets you show sources. Quality depends on good chunking, good retrieval and instructions to say 'I don't know' when the documents don't contain the answer.
Examples
- Company help desk: Answers from the latest policy documents, with links.
- School notes: A study bot answers only from your class notes.
- Citations: Each answer lists the passages it used.
How it works
- Split documents into chunks and store their embeddings.
- For each question, retrieve the most relevant chunks.
- Give those chunks to the model and ask it to answer from them, citing sources.
Check your understanding
- What does the 'retrieval' in RAG do?
- Options: Finds relevant passages to give the model; Trains a new model; Deletes old documents.
Answer: Finds relevant passages to give the model. Retrieved passages ground the answer in your documents. - Why is RAG often better than fine-tuning for company facts?
- Options: It uses up-to-date documents without retraining and can cite sources; It's always slower; It needs no documents.
Answer: It uses up-to-date documents without retraining and can cite sources. Update the documents and the answers update too.
Remember
RAG retrieves relevant passages and asks the model to answer from them, with sources.
Talk about it
What documents would you give a RAG bot for your school or office?
Go deeper
Improvements include hybrid search (keywords plus embeddings), re-ranking, query rewriting and evaluation of retrieval recall. Long-context models complement but don't remove the need for retrieval.