Stage 07 of 11 · RAG & Knowledge · 3 min · Reviewed Aug 2026
RAG and Knowledge: Chunking, Embeddings, Rerank, Alternatives
RAG is retrieval-augmented generation: fetch relevant text, then generate with that text in context. Chunking, embeddings, vector indexes, reranking, and agentic retrieval are the knobs. Graph RAG, SQL, and keyword search are alternatives — not heresies.
- Vector DBs
- Chunking
- Embeddings
- Reranking
- Alternatives
What you learn
How agents work with private knowledge without pretending the model memorized your corpus.
Why RAG exists
Weights are frozen. Your tickets, contracts, and runbooks are not in them. Fine-tuning is the wrong first tool for facts that change weekly. Retrieval puts the right passages into context so the model can cite and stay current.
If the agent answers from parametric memory on a private question, you have a hallucination with confidence. Require retrieval or an honest “not in corpus.”
Visual
RAG: Retrieve then Generate
The model never answers from memory alone. Retrieval fetches relevant passages, the context window carries them, and the answer includes citations.
Chunking is the first accuracy bug
Chunk too small and you lose the heading that makes a paragraph mean something. Chunk too large and you retrieve noise and blow the window. Prefer structure-aware splits: headings, functions, ticket comments — then a token cap with overlap.
Store metadata: source URI, updated-at, ACL, product, version. Retrieval without ACL is a leak. Retrieval without recency is stale policy.
{
"id": "runbook:pager-42#chunk-3",
"uri": "https://wiki.internal/pager-42",
"acl": ["oncall"],
"updated_at": "2026-08-01",
"tokens": 380
}Interactive
Document Ingestion Pipeline
Load document
Raw document (PDF, wiki page, ticket) enters the pipeline.