Tools & retrieval
What is RAG?
Retrieval-augmented generation fetches passages and then writes with them in context.
RAG is not automatically an agent. It becomes agentic when retrieval is a tool in a loop with state and stop conditions.
Before answering a policy question, the agent retrieves three handbook passages, checks them for relevance, and cites specific sections in the reply. The retrieval makes the answer grounded; the citation makes it verifiable.
Visual
RAG pipeline
Query → retrieve passages → augment context → generate answer.
Why it matters
Retrieval-augmented generation (RAG) solves the knowledge cutoff problem. Instead of relying on what the model memorized during training, RAG fetches current, specific information at query time and places it in context.
RAG is not automatically an agent. It becomes agentic when retrieval is a tool in a loop: search, read, decide if more retrieval is needed, verify, and answer.
How RAG works
The RAG pipeline has three stages. First, the query is used to retrieve candidate passages from a vector store, keyword index, or hybrid system. Second, the retrieved passages are placed in the model's context alongside the original question. Third, the model generates an answer using the retrieved information.
The quality of RAG depends on the retrieval step. Bad retrieval → wrong context → wrong answer, regardless of how good the model is.
RAG vs. fine-tuning
RAG and fine-tuning serve different purposes. Fine-tuning changes how the model reasons and writes. RAG changes what information is available. For dynamic knowledge (docs, policies, tickets), RAG is usually better because it does not require retraining.
Many production systems combine both: fine-tuning for tone and reasoning patterns, RAG for current facts.
Key takeaways
- 1RAG = retrieve + augment context + generate. Quality depends on retrieval.
- 2RAG becomes agentic when retrieval is a tool in a loop with state.
- 3Bad retrieval → wrong context → wrong answer, regardless of model quality.
- 4Combine RAG for facts with fine-tuning for reasoning patterns.
Common mistakes
- ✕Assuming RAG is an agent — without a tool loop and stop condition, it is a pipeline.
- ✕Indexing documents without chunking or metadata, leading to irrelevant retrieval.
- ✕Not validating retrieved passages before placing them in context.
In practice
Production RAG systems typically retrieve 3–10 passages, re-rank them by relevance, and truncate to fit the context window. The ranking step is often more important than the embedding model. Start with a simple top-k retrieval and add re-ranking when you have evaluation data.
Related terms
Concept neighborhood
Terms linked from RAG in the glossary graph.
FAQ
- Is RAG better than fine-tuning?
- They serve different purposes. RAG provides current facts; fine-tuning changes reasoning patterns. Most production systems use RAG for knowledge and fine-tuning (or good prompts) for behavior.
- How many passages should I retrieve?
- Start with 3–5. More passages add noise and consume context. Use re-ranking to ensure the top results are relevant.