Retrieval-Augmented Generation (RAG)
Retrieval-Augmented Generation (RAG) is a technique where an AI model retrieves relevant documents at query time and uses them to ground its generated answer.
RAG is why current, well-structured web content matters for AI visibility: when an engine retrieves passages to answer a question, the sources it pulls are the ones it can cite. Content written answer-first and marked up clearly is easier to retrieve accurately.
RAG is one part of how modern AI systems work. For the bigger picture — how large language models are built and how retrieval fits in — read What Are Large Language Models (LLMs)?
How it works
Source documents are split into passages, and each passage is turned into an embedding and stored in a search index. When a question arrives, it is embedded the same way, the most relevant passages are retrieved, often combined with keyword search, and those passages are placed in the model's context with an instruction to answer from them and cite them.
The answer can only be as good as retrieval: if the right passage is not found, the model either says it does not know or, worse, fills the gap from its own training.
Example
A company can let staff ask questions of its internal policies. Asked "how many days of leave can I carry over?", the system retrieves the relevant section of the HR handbook and answers with a link to it, instead of relying on what a model might guess from policies at other companies.
Common mistakes
- Chunking documents so that a question's answer is split across two passages and neither is retrieved.
- Skipping evaluation, so nobody knows how often the right passage is actually found.
- Letting the model answer when retrieval found nothing relevant, instead of saying so.