Hindsight Memory Provider
Hindsight is Kevin's leading evaluated runtime-memory candidate, not the canonical brain and not an already-active dependency on this machine.
Routing Summary
Use Hindsight for scoped cross-session runtime memory after its admission gate
passes. Route ordinary evidence lookup through bounded recall; route
multi-step memory-grounded synthesis through explicit reflect with citations
and a cost/latency receipt. Keep immutable sources and canonical facts, rules,
skills, workflows, and decisions in Kevin-Wiki/qmd. Do not call Hindsight active
until pinned install, model/data-path, bank, doctor, leakage, usefulness, and
fresh-target recovery proof exists.
Current Decision
Use Hindsight when a persistent agent needs cross-session facts, entity and temporal relationships, hybrid recall, or deliberate synthesis across prior experience. Keep QMD - Local Wiki Search Engine and the compiled wiki as the active durable truth. Hindsight may recall and derive; it does not silently promote a memory into a wiki fact, rule, skill, workflow, or decision.
As checked on 2026-08-11, the local machine has uvx and Rust but no
hindsight-embed executable, Hermes executable, coding-agent config, or proven
healthy bank. The architecture therefore says selected candidate/default
after proof, not “currently active.” No package was installed during review.
Current Source Snapshot
The MIT repository was frozen at commit
d7c33fdeafda38bb378ac869268f3b3fd53c61ea on the day v0.9.0 was current.
The core API and embedded package report version 0.9.0; the multi-harness
coding-agent integration reports 0.2.0. The captured repository contains
3,944 tracked files and current docs, tests, SDKs, MCP tools, transfer paths,
benchmarks, and integrations. Source: capture manifest, 2026-08-11
Three Operations, Three Costs
| Operation | What it owns | Kevin route |
|---|---|---|
| Retain | Ingest content and extract facts, entities, relationships, and time signals through an LLM-backed pipeline. | Feed rich, source-identified content after the immutable source revision already exists elsewhere. Use stable document IDs and actual timestamps. |
| Recall | Retrieve ranked memory with semantic, BM25, graph, and temporal strategies, followed by fusion/reranking. | First runtime-memory read when the agent should reason over evidence itself. Keep results bounded and return provenance. |
| Reflect | Run an agentic search/synthesis loop over mental models, observations, and facts, shaped by bank directives/disposition. | Use only for an explicit “why/what follows?” synthesis where added latency and inference cost are justified. Return citations and record failure/degraded mode. |
Recall returns material; reflect returns an answer. The harness must never hide that distinction by auto-reflecting on every turn without a measured benefit and a visible cost/latency receipt.
Derived Memory Hierarchy
Facts, observations, and mental models are all derived. Mental models are operator-selected standing answers that refresh as their scoped evidence changes; they can move repeated synthesis off the request path, but do not replace raw evidence or the governed owner. This is why Kevin-Wiki preserves source bytes and review receipts independently of Hindsight.
Local Is A Data Path, Not A Magic Property
Hindsight can run through a local Docker service, an embedded/local daemon, a self-hosted server, or Hindsight Cloud. “Local” is true only when every selected component stays local:
| Layer | Must be declared |
|---|---|
| API/database | local daemon, self-hosted endpoint, or cloud |
| Extraction and reflect model | Ollama/LM Studio/local model, CLI fallback, or hosted provider |
| Embeddings and reranker | exact local or remote model and version |
| Harness integration | installed targets, hook/MCP authority, config path, failure behavior |
Ollama or LM Studio can remove hosted LLM keys and provider egress. Retain, consolidation, and reflect still perform inference. A loopback API paired with a hosted provider is not an all-local memory path.
Banks, Tags, And Provenance
A bank is the hard isolation unit. Prefer one bank per user, agent authority, or repository when privacy and simple reasoning matter. Tags provide visibility inside a shared bank but are policy, not magic isolation.
For coding agents, the reviewed integration defaults to a worktree-aware bank
per repository (coding-agent::{gitProject}) shared across supported harnesses.
That is safer than one undifferentiated personal bank. If several repositories
deliberately converge into one bank:
- stamp repository, harness, session, and source identity at retain time;
- use strict tag matching for every private or user-partitioned read;
- retain an explicit bank map outside the runtime;
- test cross-repo and cross-user leakage before promotion; and
- keep source revisions in Kevin-Wiki even when Hindsight stores derived facts.
Metadata is for attribution and linking; tags are the filter surface. Never use
tags_match="any" as a user-isolation policy.
Portability And Recovery
Hindsight has meaningful export paths, but portability is an observed round trip, not the presence of an endpoint.
- Document transfer can carry documents, chunks, and extracted facts without embeddings or database IDs.
- Whole-bank transfer can additionally carry observations, bank config, mental models, directives, and webhooks.
- A fresh target re-embeds facts and reconstructs entities, links, and indexes.
- Imported observations are not automatically merged or deduplicated against an existing target; use a fresh bank or regenerate them from facts.
The adoption gate is: export a representative bank, restore into a fresh target with no ambient state, run frozen recall/reflect queries, verify provenance and scope, compare answers and cost, then prove the old provider can disappear without losing canonical source evidence.
Benchmark Boundary
The saved post carefully says “per the project's own published numbers.” Keep that qualifier. The current repository reports strong self-published results, including 94.6 on LongMemEval for its named v0.4.19 single-query setup, while the original Hindsight paper, later repo results, competing papers, and newer agent-memory benchmarks use different models, preprocessing, scales, modes, and scorers.
Use Agent Self-Improvement Eval Library to freeze corpus, memory version, answer model, extraction/preprocessing, query mode, budget, latency, cost, and scorer. LongMemEval remains useful for conversational extraction, temporal reasoning, updates, and abstention. LongMemEval-V2 adds the workflow, dynamic state, environment-gotcha, and premise-awareness cases that are closer to a coding/operator harness. Competing BEAM/LongMemEval results are counterevidence to a permanent universal-winner claim, not matched disproof unless the complete protocol is the same. [Sources: LongMemEval, Hindsight, LongMemEval-V2, and Storage Is Not Memory]
Adoption Gate
Do not call Hindsight operational until all of these are true:
- pin the core and integration versions and record the install method;
- declare local/cloud inference, embedding, reranking, and data-egress paths;
- define per-repo/user banks and any intentional shared-bank tag policy;
- wire only the selected harnesses with least authority and reversible config;
- pass readiness/doctor and retain → recall → reflect smoke tests;
- pass strict cross-bank and cross-tag leakage tests;
- compare recall-only, reflect, and no-memory baselines under one fixture;
- export and restore a representative bank into a fresh target; and
- prove a resulting correction can reach its canonical wiki/skill/workflow owner through a reviewed writeback rather than staying trapped in memory.
Boundary With Honcho
Honcho remains the peer/user/profile-modeling candidate. Hermes external memory providers are single-select, so do not run Hindsight and Honcho as two silent authorities for the same identity. If both are needed, keep Hindsight in the runtime-memory slot after proof and attach Honcho through an explicit profile/MCP/SDK boundary with a named writeback owner.
Do Not Use For
- Replacing qmd or the wiki's source-backed compiled truth.
- Storing the only copy of a source, instruction, decision, or procedure.
- Treating extracted facts or mental models as evidence without provenance.
- Calling a loopback server “fully local” while selected models use a cloud API.
- Sharing a bank across users or repositories without strict scope and leakage proof.
- Enabling auto-reflect everywhere without cost, latency, and usefulness data.
- Claiming benchmark leadership across unmatched protocols.
Timeline
- 2026-08-11 | Force-replayed the saved X post, recovered and reviewed its 72-second README video, froze Hindsight at
d7c33fde/ v0.9.0, inspected current core/docs/transfer/benchmark/coding-agent integration, and compared the benchmark claim with LongMemEval, Hindsight's paper, LongMemEval-V2, and retrieval-centered counterevidence. Corrected the local-inference, bank-isolation, recall-versus-reflect, portability, benchmark, and current-install boundaries; retained Hindsight as the leading evaluated runtime-memory candidate rather than claiming it is already active. Source: X/@itsharmanjot;vectorize-io/hindsight; capture manifest - 2026-07-04 | Added as the intended Hermes long-memory provider for Kevin's Mac Mini operator stack, with qmd/Kevin-Wiki retained as canonical durable memory and Honcho treated as a separate profile/modeling layer. Source: User request, 2026-07-04; Hermes memory providers docs