Retrieval Quality
Retrieval quality is the degree to which a search or RAG system returns the right context, at the right granularity, with enough provenance for the downstream answer or agent action to be correct.
Core idea
Retrieval is not "did we get semantically similar chunks?" It is "did the agent receive the context needed to answer or act correctly?"
Good retrieval optimizes six properties:
- Recall: the needed source appears in the candidate set.
- Precision: irrelevant context does not crowd out useful context.
- Granularity: the chunk is large enough to understand and small enough to rank.
- Freshness: current pages/sources outrank stale ones when facts change.
- Provenance: the answer can cite where context came from.
- Structural coverage: when relationships control the answer, relevant neighbors and paths survive candidate generation without overwhelming the token budget.
RAGAS and ARES both treat RAG evaluation as multi-dimensional rather than a single answer score. RAGAS separates retrieval context, faithful use of context, and final generation quality; ARES evaluates context relevance, answer faithfulness, and answer relevance using synthetic data plus a smaller human-annotated set. Source: RAGAS paper, https://arxiv.org/abs/2309.15217, 2026-06-19; ARES paper, https://arxiv.org/abs/2311.09476, 2026-06-19
Failure mode
Bad retrieval creates confident wrongness with citations nearby. The agent finds pages that sound related, quotes a plausible line, and misses the one source that actually controls the decision.
Common causes:
- embedding-only search misses exact identifiers;
- exact search misses paraphrases;
- stale chunks remain after docs change;
- chunks split away the heading, date, or source;
- reranker optimizes topicality, not answerability;
- retrieval returns many snippets but no surrounding page context;
- final answer is evaluated without evaluating retrieval separately.
Retrieval pipeline
Each stage has a distinct failure surface. If the query rewrite is bad, the retriever never sees the right intent. If candidates are weak, the reranker cannot recover. If assembly strips provenance, the answer cannot be trusted.
Metrics that matter
| Metric | What it catches |
|---|---|
| Hit rate / recall@k | Was the needed source retrieved at all? |
| MRR / nDCG | Did the needed source appear near the top? |
| Context relevance | Are retrieved passages useful for the query? |
| Context sufficiency | Is there enough context to answer without guessing? |
| Faithfulness | Does the answer stick to retrieved evidence? |
| Citation coverage | Do major claims point to sources? |
| Freshness correctness | Does the retrieval prefer newer compiled truth when appropriate? |
| Token efficiency | How much irrelevant context was stuffed into the model? |
| Path/neighbor recall | Did graph expansion recover the relationship or path needed to answer? |
| Evidence-set recall | Did the result set assemble every controlling source, not merely one plausible hit? |
| Distance robustness | Does recall survive when support is older, cross-session, or separated by distractors? |
| Rendered-state fidelity | Did the captured viewport, auth state, waits, actions, locale, and source revision expose the intended visual evidence? |
| Region grounding | Can the answer cite the exact page/tile/region that contains the visible support? |
| Cross-modal disagreement | When text, DOM, graph, and pixel lanes disagree, is the conflict surfaced and resolved instead of silently fused? |
No single metric is enough. A high-recall system can still waste the context window; a precise system can still miss the controlling document.
Memory Retrieval Profile
Memory retrieval adds time and revision to ordinary RAG evaluation. Test the same questions across short and long evidence-distance bins, before and after fact corrections, and at several retrieval budgets. Score early localization and complete evidence assembly separately. A strong top-1 hit is not sufficient when the answer depends on multiple dates, decisions, or relationship edges.
The MemoryData paper reports exactly this divergence: some compressed methods
surface one relevant item early, while linked or hierarchical methods recover
more complete support as k and evidence distance grow. It also finds that
backbone upgrades change answer quality more than they change which memory
pipeline handles updates correctly. Retrieval and temporal grounding therefore
need independent proof before generation. Source: arXiv 2606.24775v1, RQ2–RQ4,
visually verified 2026-08-10
Kevin wiki version
The wiki uses Knowledge Compilation to improve retrieval before search even starts. Instead of only retrieving raw source chunks, agents retrieve compiled pages whose current truth has already absorbed prior sources.
Retrieval quality here means:
qmd searchfinds the canonical page by slug, alias, and common phrasing;_index.mdroutes the agent to the right section;related:and inline wikilinks expose adjacent pages;- timelines preserve dated evidence;
build-indexand qmd embedding run after edits;- high-traffic routers prevent orphan concepts;
- stale pages are rewritten, not merely appended.
This is why the wiki has both RAG System Architecture and this page. Architecture names the pipeline; retrieval quality names the evaluation target.
Choosing retrieval strategy
| Need | Strategy |
|---|---|
| exact file/tool/person name | keyword/slug search first |
| fuzzy concept question | embedding or hybrid search |
| answer depends on rendered layout, tables, charts, dashboards, PDFs, or diagrams | add visual retrieval over screenshot tiles (PixelRAG) and compare/fuse with text or DOM evidence |
| local PDF/document answer | parse once with liteparse, search bounded text/BM25 windows, then screenshot exact pages only when text extraction fails |
| current external fact | web/official docs, then compile into wiki |
| high-stakes answer | multiple sources + citation check + retrieval eval |
| long research task | query decomposition + reranking + synthesis page |
| recurring query | promote retrieved synthesis into a durable concept/router |
Query-adapter admission gate
A reversible query-side embedding adapter is a candidate only after error analysis shows a stable mismatch between how users ask and how the fixed corpus is embedded. It is not the first response to stale content, missing aliases, wrong chunking, ACL bugs, a weak labeled pool, or a bad reranker.
Use this experiment contract:
- Freeze corpus and chunk revisions, base embedder, query set, relevance labels,
train/validation/test split,
k, metrics, runtime, and budget. - Keep an identity adapter as the incumbent. Build training triplets without leaking validation or test queries; include hard, semantic-opposite, random, and known false-positive negatives where appropriate.
- Compare exact/lexical, base semantic, adapted semantic, and existing hybrid retrieval on the same questions. Report precision/recall, MRR, nDCG, answer-evidence recall, per-family regressions, latency, and storage.
- Promote only strict held-out improvement with no protected-query, exact-name, freshness, ACL, or provenance regression. A corpus-average gain cannot hide a failed query family.
- Preserve the base path and adapter digest so rollback is one routing change; re-evaluate whenever the corpus, embedder, labels, or query distribution materially changes.
SantanderAI/linear-adapter-trainer v0.1.1 is a useful implementation reference:
it leaves the corpus index unchanged, initializes from identity, supports
hard/opposite/random negative mining, separates train and validation data, and
reports precision/recall/MRR/nDCG. Its template/hash demo is not admission proof
for qmd; Kevin needs a frozen benchmark built from real retrieval misses first.
Source: SantanderAI/linear-adapter-trainer@29c8b22,
reviewed 2026-08-11
Moving context and structural retrieval
“Infinite context” is useful only as shorthand. A model still receives a finite packet. The scalable design is a moving focus: load a narrow working set, notice candidate signals at its boundary, retrieve the underlying evidence or linked neighbors, and shift focus as the task develops. The evaluation question is not “can the system address every row?” but “did it recover the controlling evidence at the right moment, within budget, with authority and provenance intact?” Source: X/@daleverett X Article, 2026-07-12
Use a structural retriever only when edges contribute distinct evidence. Exact SQL remains better for predicates, joins, aggregates, and known IDs; lexical and semantic search remain better for names and paraphrases. Graph retrieval is appropriate for bounded neighborhoods, dependency chains, relationship paths, and multi-hop evidence. Evaluate it separately with path/neighbor recall, fan-out, latency, freshness after mutations, diversity, token cost, and authorization leakage tests.
pgGraph is a credible candidate for this narrow job: it keeps Postgres tables
authoritative and builds a derived CSR projection for repeated traversal. The
larger Polygres stack adds pgContext search and hybrid fusion. Neither should be
installed for Kevin Wiki merely because the article calls the result an
extended context window. Current wiki truth is Markdown and qmd already supplies
lexical/vector/reranked retrieval. A future Postgres evidence graph must beat
that baseline on representative questions and prove rebuild, freshness,
resource, tenant, and row-visibility behavior. pgGraph's own 1.0 known issues
say topology reads do not re-evaluate RLS, so it is specifically unsafe as an
unexamined private/multi-tenant graph candidate. Source: Evokoa/pgGraph at
6fd7da9; Evokoa/pgContext at 70ee2ef; Polygres site and SDK, reviewed
2026-08-10
Visual retrieval
PixelRAG adds a separate retrieval mode: render the source, index screenshot tiles, and let a vision-language model read the retrieved pixels. This improves answerability when the relevant information lives in layout, tables, charts, diagrams, dashboards, or PDF pages that text parsers flatten or drop. PixelRAG's controlled paper reports higher QA accuracy than two text-extraction baselines on six benchmarks with a current VLM reader. The result proves that a visual lane can recover otherwise lost evidence; it does not prove that pixels should replace text, DOM/accessibility, exact lookup, or graph retrieval. The paper explicitly names hybrid text/image retrieval as future work. Source: PixelRAG paper §5 and Appendix E; current source replay, 2026-08-11
The evaluation boundary is a render-state manifest, not merely an image. If the browser cannot see content because it is behind login, hidden in a collapsed section, blocked by a captcha/paywall, still loading, localized differently, or using a different source revision, screenshot retrieval will faithfully index the wrong state. Each visual evidence item should therefore preserve:
- canonical URL and immutable source/content revision when obtainable;
- capture timestamp, viewport and device scale, locale/theme, authentication scope, and personalization boundary;
- waits, clicks, scrolls, tab/menu expansion, consent handling, and blocked-resource status;
- page/tile/region coordinates plus a content hash;
- hyperlink targets and anchor text as side metadata, because pixels render links but cannot follow them;
- moderation, privacy, and rights decisions for the image itself.
Version 0.4's source-manifest and atomic-write changes are the right implementation pattern: one stable source identity owns the rendered artifacts, partial reruns do not silently corrupt the corpus, and waits are bounded. Kevin-Wiki should generalize that contract across browser and document capture rather than depend on PixelRAG-specific storage. Source: PixelRAG v0.4.0 release and source tests, reviewed 2026-08-11
Use a dual-channel evaluation for representative queries. Measure text/DOM, visual, and fused lanes separately under the same reader, k, corpus revision, and budget. Score answer/evidence recall, region grounding, faithfulness, disagreement, latency, tokens, storage, rights/privacy exposure, and failure under auth/timing changes. PixelRAG's full paper reproduction is expensive—an H100 reader, multiple GPU services, hundreds of gigabytes of indexes, and potentially terabytes of tiles—so a bounded local corpus should prove the route before any large download or production service adoption. Source: PixelRAG eval/README.md, reviewed 2026-08-11
Retrieval backend is a workload profile
The saved “turbopuffer is the goat” claim is useful because it names real
alternatives, but it does not define a winner without a workload. Compare the
complete system under the same source revisions, chunking, embeddings, hybrid
query, filters/ACLs, freshness target, k, reranker, answer model, load, region,
and failure/recovery procedure.
| Route | Best first fit | Required proof |
|---|---|---|
| qmd + canonical Markdown | Kevin's local compiled knowledge, exact/semantic retrieval, portable source ownership | real query set, freshness and alias coverage, citation/source recovery, local latency and rebuild behavior |
| PostgreSQL + pgvector | app already owns relational truth, transactions, RLS/joins, moderate vector scale, and database operations | exact vs HNSW/IVFFlat recall, filtered-query behavior, index build/memory, vacuum/backup/restore, tenant/RLS and migration |
| Turbopuffer | serverless vector + full-text/hybrid search where object-storage economics and low operations are measured advantages | hot/cold namespace latency, ingestion/freshness, consistency/delete/export, filter/ACL behavior, region, egress, availability and dated cost |
| OpenSearch | high-throughput hybrid search, analytics, aggregations/facets, managed AWS/IAM integration | cluster/serverless operations, relevance tuning, IAM/tenant policy, shard/index lifecycle, latency, recovery and full cost |
| S3 Vectors / OpenSearch S3 engine | very large or colder vector corpora with higher latency tolerance and cost pressure | current feature limits, sub-second target, KMS/IAM, update/delete semantics, export or engine synchronization, duplicated-storage cost |
| AgentCore Memory | managed short/long conversation memory with actor/session/namespace and extraction strategies | event expiry, actor isolation, strategy quality, correction/deletion/export, KMS/region, provider cost and portability; it is not merely a vector database |
Turbopuffer's current site reports object-storage-backed serverless search, hybrid/full-text support and vendor benchmark figures; pgvector documents exact and approximate transactional Postgres search plus filtering caveats; AWS documents one-time S3-vector exports, a cost-oriented S3 engine, and AgentCore's event/strategy/actor model. These are distinct products. Promote a route only when the same representative corpus proves answer/evidence quality, freshness, latency distribution, load, operations, privacy, failure recovery, portability, and total cost. Source: official Turbopuffer docs/site; pgvector README; AWS S3 Vectors, OpenSearch, and AgentCore Memory documentation, checked 2026-08-12
Content-bearing search
Search tools should return clean source content plus title, canonical URL,
author/date, stable locator, extraction state, and failure metadata when the job
needs evidence—not only snippets that force another discovery round. Claims of
a fixed token multiplier remain benchmark leads. Compare discovery-only and
content-bearing lanes on source coverage, extraction fidelity, citations,
latency, tokens, cost, blocks, freshness, and whether the downstream agent still
opens the primary source for controlling claims. Source: X
2077753829526056985; resolved article replay, 2026-08-12
Design implications
Agent products should show retrieval evidence:
- top sources used;
- skipped near-matches when useful;
- freshness dates;
- citation confidence;
- missing-source warning;
- "open full context" affordance.
Search UX should optimize answerability, not just matching. The best result is the one that lets the agent take the next correct step.
Failure modes
- hiding retrieval behind an answer with no sources;
- ranking old docs above current compiled truth;
- using vector search for exact command names;
- using keyword search for conceptual paraphrases;
- measuring answer quality while retrieval silently fails;
- treating citations as decoration rather than proof.
Timeline
-
2026-08-12 | Replaced a universal vector-store ranking with a workload matrix across qmd, pgvector, Turbopuffer, OpenSearch/S3 Vectors, and AgentCore Memory; added a content-bearing search receipt and required same-corpus quality, freshness, latency, operations, privacy, recovery, portability, and cost proof. Source: X
2080696750856393055; official provider documentation; X2077753829526056985 -
2026-08-11 | Added a query-adapter admission gate: require real labeled qmd failures, frozen splits and base embeddings, an identity baseline, per-family retrieval/answer regressions, ACL/provenance checks, and reversible routing before trialing a linear adapter. Source: SantanderAI
linear-adapter-trainer@29c8b22; exact-source replay -
2026-08-11 | Replayed PixelRAG's complete current source, paper, evaluation package, release, thread, and video. Recast visual retrieval as one hybrid evidence lane; added render-state manifests, actionable-link side metadata, region grounding, cross-modal disagreement, moderation/rights, same-condition evaluation, and cost/reproducibility gates. Source: PixelRAG current replay, 2026-08-11
-
2026-08-10 | Added evidence-set recall, temporal-distance robustness, update-state testing, and the requirement to evaluate retrieval independently of a stronger answer model. Source: arXiv 2606.24775v1, RQ2–RQ4
-
2026-08-10 | Added moving-focus context and structural retrieval as a separately evaluated lane. Retained pgGraph/Polygres as a relationship-heavy Postgres candidate, not a qmd replacement, and made topology authorization leakage a required eval. Source: X/@daleverett X Article; pgGraph/pgContext/Polygres primary sources, reviewed 2026-08-10
-
2026-06-30 | Added visual retrieval as a first-class retrieval strategy after deep-reviewing PixelRAG. The durable distinction: screenshot-tile retrieval improves layout-heavy answerability, but browser-state setup remains part of retrieval quality. Source: X bookmark artifact audit, 2026-06-30
-
2026-06-30 | Added LiteParse as the local document QA route: parse once, search bounded windows, and escalate to screenshots or visual retrieval only when text/layout extraction cannot answer. Source: X bookmark artifact audit, 2026-06-30
-
2026-06-19 | Created as the missing evaluation concept behind qmd/wiki search, RAG architecture, context engineering, and retrieval-augmented agents. Grounded against RAGAS and ARES. Source: User request; RAGAS; ARES; local qmd/RAG pages, 2026-06-19