RAG System Architecture
A RAG system is two pipelines sharing an index: offline ingestion turns sources into retrievable units; online query selects the right units for a model or agent.
The conceptual spectrum lives in Retrieval-Augmented Agents. This page describes the architecture.
The evaluation target lives in Retrieval Quality: answerability, citation coverage, freshness, context sufficiency, and faithful use of retrieved evidence. Keep the split clear. Architecture says how retrieval is assembled; quality says whether it found the right thing.
Offline Ingestion Pipeline
Steps:
- Load and parse: preserve structure, title, URL, source path, and timestamp.
- Normalize: strip boilerplate, keep headings, preserve code blocks when relevant.
- Chunk: split by semantic structure before fixed size. Overlap only where boundary loss matters.
- Embed: compute vectors for semantic recall.
- Index keywords: exact names, commands, file paths, identifiers.
- Store metadata: citations and freshness require it.
Most RAG quality problems are ingestion problems disguised as prompt problems.
Online Query Pipeline
Steps:
- Query transform: rewrite, expand, or split multi-hop tasks.
- Filter scope first: apply privacy, ACL, project/workflow, and freshness boundaries before candidate retrieval.
- Parallel retrieve: combine semantic, exact/lexical, rarity, recency, and direct live-tool evidence.
- Fuse and diversify: use weighted reciprocal-rank fusion, consolidate duplicate chunks into sources, and cap domination by any one source.
- Rerank: prioritize precision against the actual question after cheap recall.
- Expand context: add neighboring sections or the complete semantic unit around each winning hit.
- Assemble evidence: pack citations, provenance, dates, privacy, and why each item was loaded.
- Generate or act: answer, call tools, or continue agent loop.
- Write back: update memory only through the owning workflow and authority boundary.
Kevin Wiki Implementation
The wiki uses a document-first variant:
- Markdown pages are the source corpus.
build-index.tsproduces catalog/backlink/discovery artifacts.- qmd provides retrieval and embedding.
- Exact search and source-specific tools preserve identifiers and live evidence that embeddings miss.
- Weighted RRF can combine qmd, lexical, recency, and tool result lists without pretending their raw scores are comparable.
_index.mdgives a human-readable catalog.- Timelines and frontmatter preserve provenance.
qmd update && qmd embedrefreshes the retrieval layer after edits.
This is intentionally not a black-box vector DB. The corpus remains readable and editable.
Design Implications
- Keep ingestion and query concerns separate.
- Do not use embeddings alone; exact identifiers matter.
- Filter ACL/project scope before retrieval and keep source authority attached through synthesis.
- Fuse independent rank lists, then dedupe, cap source concentration, rerank, and expand context.
- Always store enough metadata to cite and refresh.
- Reranking is often cheaper than overhauling the whole embedding model.
- Evaluate retrieval recall before blaming the generator.
- Re-index after maintenance or large doc edits.
Failure Modes
- Stale embeddings after docs change.
- Chunks too small to preserve meaning.
- Chunks too large to match precise queries.
- Missing exact keyword index.
- No source metadata.
- Agent reads a retrieved snippet but not the page context.
Architecture Position
| Axis | Value |
|---|---|
| Family | Memory, context, and retrieval |
| Boundary owned | Ingestion/query split for retrieval augmented generation. |
| Read with | Vector Database and ANN Index Architecture, Retrieval Quality, Agent Memory System Architecture |
| Use this page when | designing source-backed answer systems |
Timeline
-
2026-07-16 | Added the Cerebras production retrieval reference: privacy/project filtering, parallel lexical/semantic/rarity/recency candidates, weighted RRF, source dedupe/diversity caps, question-specific reranking, semantic-unit context expansion, and rich evidence packets. Source: Cerebras Enterprise Knowledge Base Reference Architecture
-
2026-07-01 | Architecture category refresh added this page to the Memory, context, and retrieval family, linked it to Architecture System Map, and kept it standalone because it owns this boundary: Ingestion/query split for retrieval augmented generation. Source: User request, 2026-07-01
-
2026-06-18 | Expanded with Kevin's qmd/wiki implementation, ingestion/query diagrams, and failure modes. Source: User request, 2026-06-18
-
2026-06-19 | Linked Retrieval Quality as the evaluation layer for this architecture. Source: whole-wiki concept review, 2026-06-19
-
2026-06-02 | Page created. Captured offline-ingestion / online-query split, chunking as central knob, hybrid retrieval, reranking, separate retrieval/generation evaluation, and mapping to qmd/context engine. Source: compiled from RAG system-design literature