Retrieval Quality

Retrieval quality is the degree to which a search or RAG system returns the right context, at the right granularity, with enough provenance for the downstream answer or agent action to be correct.

Core idea

Retrieval is not "did we get semantically similar chunks?" It is "did the agent receive the context needed to answer or act correctly?"

Good retrieval optimizes six properties:

  1. Recall: the needed source appears in the candidate set.
  2. Precision: irrelevant context does not crowd out useful context.
  3. Granularity: the chunk is large enough to understand and small enough to rank.
  4. Freshness: current pages/sources outrank stale ones when facts change.
  5. Provenance: the answer can cite where context came from.
  6. Structural coverage: when relationships control the answer, relevant neighbors and paths survive candidate generation without overwhelming the token budget.

RAGAS and ARES both treat RAG evaluation as multi-dimensional rather than a single answer score. RAGAS separates retrieval context, faithful use of context, and final generation quality; ARES evaluates context relevance, answer faithfulness, and answer relevance using synthetic data plus a smaller human-annotated set. Source: RAGAS paper, https://arxiv.org/abs/2309.15217, 2026-06-19; ARES paper, https://arxiv.org/abs/2311.09476, 2026-06-19

Failure mode

Bad retrieval creates confident wrongness with citations nearby. The agent finds pages that sound related, quotes a plausible line, and misses the one source that actually controls the decision.

Common causes:

  • embedding-only search misses exact identifiers;
  • exact search misses paraphrases;
  • stale chunks remain after docs change;
  • chunks split away the heading, date, or source;
  • reranker optimizes topicality, not answerability;
  • retrieval returns many snippets but no surrounding page context;
  • final answer is evaluated without evaluating retrieval separately.

Retrieval pipeline

Each stage has a distinct failure surface. If the query rewrite is bad, the retriever never sees the right intent. If candidates are weak, the reranker cannot recover. If assembly strips provenance, the answer cannot be trusted.

Metrics that matter

Metric What it catches
Hit rate / recall@k Was the needed source retrieved at all?
MRR / nDCG Did the needed source appear near the top?
Context relevance Are retrieved passages useful for the query?
Context sufficiency Is there enough context to answer without guessing?
Faithfulness Does the answer stick to retrieved evidence?
Citation coverage Do major claims point to sources?
Freshness correctness Does the retrieval prefer newer compiled truth when appropriate?
Token efficiency How much irrelevant context was stuffed into the model?
Path/neighbor recall Did graph expansion recover the relationship or path needed to answer?
Evidence-set recall Did the result set assemble every controlling source, not merely one plausible hit?
Distance robustness Does recall survive when support is older, cross-session, or separated by distractors?
Rendered-state fidelity Did the captured viewport, auth state, waits, actions, locale, and source revision expose the intended visual evidence?
Region grounding Can the answer cite the exact page/tile/region that contains the visible support?
Cross-modal disagreement When text, DOM, graph, and pixel lanes disagree, is the conflict surfaced and resolved instead of silently fused?

No single metric is enough. A high-recall system can still waste the context window; a precise system can still miss the controlling document.

Memory Retrieval Profile

Memory retrieval adds time and revision to ordinary RAG evaluation. Test the same questions across short and long evidence-distance bins, before and after fact corrections, and at several retrieval budgets. Score early localization and complete evidence assembly separately. A strong top-1 hit is not sufficient when the answer depends on multiple dates, decisions, or relationship edges.

The MemoryData paper reports exactly this divergence: some compressed methods surface one relevant item early, while linked or hierarchical methods recover more complete support as k and evidence distance grow. It also finds that backbone upgrades change answer quality more than they change which memory pipeline handles updates correctly. Retrieval and temporal grounding therefore need independent proof before generation. Source: arXiv 2606.24775v1, RQ2–RQ4, visually verified 2026-08-10

Kevin wiki version

The wiki uses Knowledge Compilation to improve retrieval before search even starts. Instead of only retrieving raw source chunks, agents retrieve compiled pages whose current truth has already absorbed prior sources.

Retrieval quality here means:

  • qmd search finds the canonical page by slug, alias, and common phrasing;
  • _index.md routes the agent to the right section;
  • related: and inline wikilinks expose adjacent pages;
  • timelines preserve dated evidence;
  • build-index and qmd embedding run after edits;
  • high-traffic routers prevent orphan concepts;
  • stale pages are rewritten, not merely appended.

This is why the wiki has both RAG System Architecture and this page. Architecture names the pipeline; retrieval quality names the evaluation target.

Choosing retrieval strategy

Need Strategy
exact file/tool/person name keyword/slug search first
fuzzy concept question embedding or hybrid search
answer depends on rendered layout, tables, charts, dashboards, PDFs, or diagrams add visual retrieval over screenshot tiles (PixelRAG) and compare/fuse with text or DOM evidence
local PDF/document answer parse once with liteparse, search bounded text/BM25 windows, then screenshot exact pages only when text extraction fails
current external fact web/official docs, then compile into wiki
high-stakes answer multiple sources + citation check + retrieval eval
long research task query decomposition + reranking + synthesis page
recurring query promote retrieved synthesis into a durable concept/router

Query-adapter admission gate

A reversible query-side embedding adapter is a candidate only after error analysis shows a stable mismatch between how users ask and how the fixed corpus is embedded. It is not the first response to stale content, missing aliases, wrong chunking, ACL bugs, a weak labeled pool, or a bad reranker.

Use this experiment contract:

  1. Freeze corpus and chunk revisions, base embedder, query set, relevance labels, train/validation/test split, k, metrics, runtime, and budget.
  2. Keep an identity adapter as the incumbent. Build training triplets without leaking validation or test queries; include hard, semantic-opposite, random, and known false-positive negatives where appropriate.
  3. Compare exact/lexical, base semantic, adapted semantic, and existing hybrid retrieval on the same questions. Report precision/recall, MRR, nDCG, answer-evidence recall, per-family regressions, latency, and storage.
  4. Promote only strict held-out improvement with no protected-query, exact-name, freshness, ACL, or provenance regression. A corpus-average gain cannot hide a failed query family.
  5. Preserve the base path and adapter digest so rollback is one routing change; re-evaluate whenever the corpus, embedder, labels, or query distribution materially changes.

SantanderAI/linear-adapter-trainer v0.1.1 is a useful implementation reference: it leaves the corpus index unchanged, initializes from identity, supports hard/opposite/random negative mining, separates train and validation data, and reports precision/recall/MRR/nDCG. Its template/hash demo is not admission proof for qmd; Kevin needs a frozen benchmark built from real retrieval misses first. Source: SantanderAI/linear-adapter-trainer@29c8b22, reviewed 2026-08-11

Moving context and structural retrieval

“Infinite context” is useful only as shorthand. A model still receives a finite packet. The scalable design is a moving focus: load a narrow working set, notice candidate signals at its boundary, retrieve the underlying evidence or linked neighbors, and shift focus as the task develops. The evaluation question is not “can the system address every row?” but “did it recover the controlling evidence at the right moment, within budget, with authority and provenance intact?” Source: X/@daleverett X Article, 2026-07-12

Use a structural retriever only when edges contribute distinct evidence. Exact SQL remains better for predicates, joins, aggregates, and known IDs; lexical and semantic search remain better for names and paraphrases. Graph retrieval is appropriate for bounded neighborhoods, dependency chains, relationship paths, and multi-hop evidence. Evaluate it separately with path/neighbor recall, fan-out, latency, freshness after mutations, diversity, token cost, and authorization leakage tests.

pgGraph is a credible candidate for this narrow job: it keeps Postgres tables authoritative and builds a derived CSR projection for repeated traversal. The larger Polygres stack adds pgContext search and hybrid fusion. Neither should be installed for Kevin Wiki merely because the article calls the result an extended context window. Current wiki truth is Markdown and qmd already supplies lexical/vector/reranked retrieval. A future Postgres evidence graph must beat that baseline on representative questions and prove rebuild, freshness, resource, tenant, and row-visibility behavior. pgGraph's own 1.0 known issues say topology reads do not re-evaluate RLS, so it is specifically unsafe as an unexamined private/multi-tenant graph candidate. Source: Evokoa/pgGraph at 6fd7da9; Evokoa/pgContext at 70ee2ef; Polygres site and SDK, reviewed 2026-08-10

Visual retrieval

PixelRAG adds a separate retrieval mode: render the source, index screenshot tiles, and let a vision-language model read the retrieved pixels. This improves answerability when the relevant information lives in layout, tables, charts, diagrams, dashboards, or PDF pages that text parsers flatten or drop. PixelRAG's controlled paper reports higher QA accuracy than two text-extraction baselines on six benchmarks with a current VLM reader. The result proves that a visual lane can recover otherwise lost evidence; it does not prove that pixels should replace text, DOM/accessibility, exact lookup, or graph retrieval. The paper explicitly names hybrid text/image retrieval as future work. Source: PixelRAG paper §5 and Appendix E; current source replay, 2026-08-11

The evaluation boundary is a render-state manifest, not merely an image. If the browser cannot see content because it is behind login, hidden in a collapsed section, blocked by a captcha/paywall, still loading, localized differently, or using a different source revision, screenshot retrieval will faithfully index the wrong state. Each visual evidence item should therefore preserve:

  • canonical URL and immutable source/content revision when obtainable;
  • capture timestamp, viewport and device scale, locale/theme, authentication scope, and personalization boundary;
  • waits, clicks, scrolls, tab/menu expansion, consent handling, and blocked-resource status;
  • page/tile/region coordinates plus a content hash;
  • hyperlink targets and anchor text as side metadata, because pixels render links but cannot follow them;
  • moderation, privacy, and rights decisions for the image itself.

Version 0.4's source-manifest and atomic-write changes are the right implementation pattern: one stable source identity owns the rendered artifacts, partial reruns do not silently corrupt the corpus, and waits are bounded. Kevin-Wiki should generalize that contract across browser and document capture rather than depend on PixelRAG-specific storage. Source: PixelRAG v0.4.0 release and source tests, reviewed 2026-08-11

Use a dual-channel evaluation for representative queries. Measure text/DOM, visual, and fused lanes separately under the same reader, k, corpus revision, and budget. Score answer/evidence recall, region grounding, faithfulness, disagreement, latency, tokens, storage, rights/privacy exposure, and failure under auth/timing changes. PixelRAG's full paper reproduction is expensive—an H100 reader, multiple GPU services, hundreds of gigabytes of indexes, and potentially terabytes of tiles—so a bounded local corpus should prove the route before any large download or production service adoption. Source: PixelRAG eval/README.md, reviewed 2026-08-11

Retrieval backend is a workload profile

The saved “turbopuffer is the goat” claim is useful because it names real alternatives, but it does not define a winner without a workload. Compare the complete system under the same source revisions, chunking, embeddings, hybrid query, filters/ACLs, freshness target, k, reranker, answer model, load, region, and failure/recovery procedure.

Route Best first fit Required proof
qmd + canonical Markdown Kevin's local compiled knowledge, exact/semantic retrieval, portable source ownership real query set, freshness and alias coverage, citation/source recovery, local latency and rebuild behavior
PostgreSQL + pgvector app already owns relational truth, transactions, RLS/joins, moderate vector scale, and database operations exact vs HNSW/IVFFlat recall, filtered-query behavior, index build/memory, vacuum/backup/restore, tenant/RLS and migration
Turbopuffer serverless vector + full-text/hybrid search where object-storage economics and low operations are measured advantages hot/cold namespace latency, ingestion/freshness, consistency/delete/export, filter/ACL behavior, region, egress, availability and dated cost
OpenSearch high-throughput hybrid search, analytics, aggregations/facets, managed AWS/IAM integration cluster/serverless operations, relevance tuning, IAM/tenant policy, shard/index lifecycle, latency, recovery and full cost
S3 Vectors / OpenSearch S3 engine very large or colder vector corpora with higher latency tolerance and cost pressure current feature limits, sub-second target, KMS/IAM, update/delete semantics, export or engine synchronization, duplicated-storage cost
AgentCore Memory managed short/long conversation memory with actor/session/namespace and extraction strategies event expiry, actor isolation, strategy quality, correction/deletion/export, KMS/region, provider cost and portability; it is not merely a vector database

Turbopuffer's current site reports object-storage-backed serverless search, hybrid/full-text support and vendor benchmark figures; pgvector documents exact and approximate transactional Postgres search plus filtering caveats; AWS documents one-time S3-vector exports, a cost-oriented S3 engine, and AgentCore's event/strategy/actor model. These are distinct products. Promote a route only when the same representative corpus proves answer/evidence quality, freshness, latency distribution, load, operations, privacy, failure recovery, portability, and total cost. Source: official Turbopuffer docs/site; pgvector README; AWS S3 Vectors, OpenSearch, and AgentCore Memory documentation, checked 2026-08-12

Search tools should return clean source content plus title, canonical URL, author/date, stable locator, extraction state, and failure metadata when the job needs evidence—not only snippets that force another discovery round. Claims of a fixed token multiplier remain benchmark leads. Compare discovery-only and content-bearing lanes on source coverage, extraction fidelity, citations, latency, tokens, cost, blocks, freshness, and whether the downstream agent still opens the primary source for controlling claims. Source: X 2077753829526056985; resolved article replay, 2026-08-12

Design implications

Agent products should show retrieval evidence:

  • top sources used;
  • skipped near-matches when useful;
  • freshness dates;
  • citation confidence;
  • missing-source warning;
  • "open full context" affordance.

Search UX should optimize answerability, not just matching. The best result is the one that lets the agent take the next correct step.

Failure modes

  • hiding retrieval behind an answer with no sources;
  • ranking old docs above current compiled truth;
  • using vector search for exact command names;
  • using keyword search for conceptual paraphrases;
  • measuring answer quality while retrieval silently fails;
  • treating citations as decoration rather than proof.

Timeline