RAG System Architecture

A RAG system is two pipelines sharing an index: offline ingestion turns sources into retrievable units; online query selects the right units for a model or agent.

The conceptual spectrum lives in Retrieval-Augmented Agents. This page describes the architecture.

The evaluation target lives in Retrieval Quality: answerability, citation coverage, freshness, context sufficiency, and faithful use of retrieved evidence. Keep the split clear. Architecture says how retrieval is assembled; quality says whether it found the right thing.

Offline Ingestion Pipeline

Steps:

  1. Load and parse: preserve structure, title, URL, source path, and timestamp.
  2. Normalize: strip boilerplate, keep headings, preserve code blocks when relevant.
  3. Chunk: split by semantic structure before fixed size. Overlap only where boundary loss matters.
  4. Embed: compute vectors for semantic recall.
  5. Index keywords: exact names, commands, file paths, identifiers.
  6. Store metadata: citations and freshness require it.

Most RAG quality problems are ingestion problems disguised as prompt problems.

Online Query Pipeline

Steps:

  1. Query transform: rewrite, expand, or split multi-hop tasks.
  2. Filter scope first: apply privacy, ACL, project/workflow, and freshness boundaries before candidate retrieval.
  3. Parallel retrieve: combine semantic, exact/lexical, rarity, recency, and direct live-tool evidence.
  4. Fuse and diversify: use weighted reciprocal-rank fusion, consolidate duplicate chunks into sources, and cap domination by any one source.
  5. Rerank: prioritize precision against the actual question after cheap recall.
  6. Expand context: add neighboring sections or the complete semantic unit around each winning hit.
  7. Assemble evidence: pack citations, provenance, dates, privacy, and why each item was loaded.
  8. Generate or act: answer, call tools, or continue agent loop.
  9. Write back: update memory only through the owning workflow and authority boundary.

Kevin Wiki Implementation

The wiki uses a document-first variant:

  • Markdown pages are the source corpus.
  • build-index.ts produces catalog/backlink/discovery artifacts.
  • qmd provides retrieval and embedding.
  • Exact search and source-specific tools preserve identifiers and live evidence that embeddings miss.
  • Weighted RRF can combine qmd, lexical, recency, and tool result lists without pretending their raw scores are comparable.
  • _index.md gives a human-readable catalog.
  • Timelines and frontmatter preserve provenance.
  • qmd update && qmd embed refreshes the retrieval layer after edits.

This is intentionally not a black-box vector DB. The corpus remains readable and editable.

Design Implications

  • Keep ingestion and query concerns separate.
  • Do not use embeddings alone; exact identifiers matter.
  • Filter ACL/project scope before retrieval and keep source authority attached through synthesis.
  • Fuse independent rank lists, then dedupe, cap source concentration, rerank, and expand context.
  • Always store enough metadata to cite and refresh.
  • Reranking is often cheaper than overhauling the whole embedding model.
  • Evaluate retrieval recall before blaming the generator.
  • Re-index after maintenance or large doc edits.

Failure Modes

  • Stale embeddings after docs change.
  • Chunks too small to preserve meaning.
  • Chunks too large to match precise queries.
  • Missing exact keyword index.
  • No source metadata.
  • Agent reads a retrieved snippet but not the page context.

Architecture Position

Axis Value
Family Memory, context, and retrieval
Boundary owned Ingestion/query split for retrieval augmented generation.
Read with Vector Database and ANN Index Architecture, Retrieval Quality, Agent Memory System Architecture
Use this page when designing source-backed answer systems

Timeline