Agent Memory System Architecture
The memory subsystem moves facts between volatile working context and durable long-term stores. Its quality is determined by extraction, consolidation, retrieval, and conflict handling.
The cognitive taxonomy lives in Agent Memory Patterns. This page is the system design.
Kevin Retrieval Stack
Kevin's current memory/retrieval stack has four distinct lanes:
| Lane | Owner | Use |
|---|---|---|
| Compiled durable memory | Kevin-Wiki + QMD - Local Wiki Search Engine | People, tools, decisions, workflows, source-backed synthesis, and operational rules. |
| Runtime long memory | Hindsight Memory Provider candidate inside Hermes Agent (Nous Research) | Cross-session facts, hybrid recall, entity/temporal relationships, and explicit synthesis after install, scope, doctor, leakage, usefulness, and replay proof. |
| Profile/user model | Honcho | Kevin/preferences/project/agent peer modeling, attached through MCP/SDK or an explicit bridge. |
| Repo topology | Graphify through Agent-Docs Mesh repo-graph blocks |
Path/explain/affected questions, PR risk, and unfamiliar codebase orientation. |
The boundary matters. qmd answers "what is the compiled truth?" A proven
Hindsight deployment answers "what did this scoped runtime retain and retrieve
across runs?" Honcho answers "what model of Kevin or a peer is useful here?"
Graphify answers "what touches what in this repo?" None of the runtime lanes
replace the wiki. A runtime memory becomes durable only after it is written back
to a page, skill, automation, config file, project AGENTS.md, or log. At the
2026-08-11 review point, Hindsight is selected and source-evaluated but not
installed or healthy on this machine; do not collapse architecture intent into
deployment state.
Evidence-To-Context Layers
The raw layer is lossless and content-addressed. The stable-object layer is selective and authoritative. The projection layer is rebuildable. The working set is finite and disposable. This makes “infinite context” an addressability property of the store, not a claim about prompt size. A failed promotion can be replayed from the source; a stale index can be rebuilt from the object; a bad answer does not rewrite memory without a new evidence-bound decision.
Temporal Signal Envelope
Capture time and truth time are different. A source can be fetched today while describing a fact that was true last year; a git commit records when text changed, not when the claim became or stopped being true. Time-sensitive atomic signals therefore support this backward-compatible envelope:
temporal?: {
status: "current" | "historical" | "superseded" | "uncertain";
recordedAt: string;
validFrom?: string;
validTo?: string;
supersedes?: string[];
}
capturedAt remains source-transaction time. recordedAt is when the
system accepted the signal representation. validFrom and validTo bound
real-world validity when evidence supports them. supersedes creates an
auditable correction edge; it never deletes the prior signal or source. Missing
valid-world dates remain unknown rather than inferred from capture time.
scripts/lib/signal-review.ts validates the envelope while older reviews
remain valid. [Sources: Zep fact
metadata; X/@michael_chomsky, reviewed
2026-08-11]
Compaction Is A Durable Write Boundary
Compaction is a lossy working-context operation, not the durable memory store. Before a harness discards volatile context, it should persist one idempotent checkpoint containing:
- the exact source/event watermark and checkpoint ID;
- pending decisions, open questions, approvals, blockers, and tool state;
- newly proposed memories with provenance and authority scope;
- artifact references needed to reconstruct the run; and
- the tail boundary from which the next context resumes.
Only a successful checkpoint event may advance the watermark. Raw session messages remain recoverable where the harness supports them, and frozen retrieval/continuity checks run after resumption. OpenCode's checkpoint-plus-tail model and OpenClaw's pre-compaction memory flush support this boundary. They do not establish the source post's unsupported claim that Codex compaction is encrypted or unexportable; that claim remains outside canonical truth. [Sources: OpenCode compaction; OpenClaw memory, reviewed 2026-08-11]
One Canonical Writer Per Memory Class
A harness memory and an external provider may coexist only when ownership is explicit. For each semantic, episodic, procedural, profile, and runtime-memory class, designate one canonical writer. Other adapters are read-only consumers, mirrors, or rebuildable projections. Two peers must not silently last-write-win the same claim. Cross-store promotion requires provenance, a dedup/supersession decision, and a write receipt. This preserves portability without creating two conflicting authorities.
Import is not consent to universal recall
An importer such as Nessie can make ChatGPT, Perplexity, Gemini, Claude, or other histories portable, but transport does not erase the original audience, account, subject, retention, deletion, or confidentiality boundary. Preserve the export revision and source account; parse conversations into attributable units; classify participants and sensitive fields; deduplicate without erasing conflicts; keep imported raw history separate from derived facts/preferences; and require explicit destination scopes before indexing or exposing it through MCP. Support preview, correction, per-source exclusion, expiry, export, verified delete, and fresh-target restore. An MCP server that can recall the corpus is a new disclosure surface, not merely a storage convenience.
Personal-data removal is the inverse workflow and needs equally strong proof. Use official bulk mechanisms first; a consent flag is not signed authority; an opt-out request is not confirmed removal; and a later rescan is required because brokers relist data. Unbroker already owns the detailed assisted, encrypted, receipt-bearing route.
Portable Agent Surface
A memory system needs a small job-shaped interface even when its internal data model is richer. The portable surface should cover five jobs:
| Job | Required contract |
|---|---|
| Recall | Bounded results, evidence for why each result matched, source and authority scope, honest degraded mode, and token-budget accounting |
| Remember | One typed claim or event per write, mandatory provenance, entity/object scope, visibility, dedup/supersession result, and optional expiry |
| Resolve owner/entity | Deterministic alias/title/ID resolution, bounded suggestions, no error on an ordinary miss, and no private-field leakage |
| Synthesize | Explicitly expensive evidence combination with sources, gaps, model/cost receipt, and an unavailable state instead of a fabricated answer |
| Forget/retire | Idempotent expiry or retirement with reason and audit history; immutable source evidence is never destructively erased |
GBrain's frozen five verbs are one concrete implementation of these job
classes, not an automatic standard for Kevin-Wiki. Kevin's system preserves the
evidence/object distinction: remember cannot turn an unsupported sentence
directly into compiled truth, and forget cannot delete a source revision or
review receipt.
Kevin-Wiki now implements the first portable write slice in
@kevin-wiki/brain-runtime: remember, scoped recall, and retire over a
private hash-chained JSONL ledger. POST /api/memory requires a bearer token and
an Idempotency-Key; the server derives the authority identity, while the
caller declares an explicit source scope. A write requires one typed claim,
visibility, at least one existing brain/objects.json owner, and provenance
containing source ID, revision ID, storage reference, and SHA-256 content hash.
The resulting state is proposed, never compiled truth. Retirement hides the
proposal from ordinary recall while preserving the record and evidence in
history. Conflicting idempotency reuse, unknown owners, malformed provenance,
scope mismatches, and corrupt ledger chains fail closed.
This adapter is implemented and installed as a persistent single-writer macOS
LaunchAgent inside briefd. Its mode-0600 token is absent from the plist, the
service binds only to loopback, authority is derived by the server, the default
source allowlist is exactly repo:kevin-wiki, and public memory is denied by
default. MemoryHttpClient gives any local Codex, Claude Code, Cursor, Hermes,
or TypeScript adapter the same API without repository internals. Every accepted
remember or retire event opens a private proposal-only Daily Brief card; it
does not mutate canonical truth. The ledger fsyncs each digest-chained event
before acknowledgement.
The repeatable npm run memory:conformance gate proves authenticated
write/read/retire, idempotent replay, source-isolation denial, public-write
denial, retained history, daemon restart durability, and Daily Brief
projection against the installed service. The alternate Next.js adapter fails
closed in production without a token, explicit source scopes, and a durable
ledger. Multi-instance coordination, portable synthesis, richer owner
resolution, pushed context, and the cross-harness usefulness benchmark remain
separate future capabilities rather than blockers for the portable write
contract. Source: packages/brain-runtime/src/memory.ts;
packages/brain-runtime/src/memory-client.ts; apps/briefd/src/server.ts;
scripts/memory-conformance.ts; installed LaunchAgent replay, 2026-08-12
Every operation resolves two independent scopes before reading or writing:
- authority/brain scope — who owns the store and policy; and
- source/repository scope — which corpus inside that authority boundary.
Cross-authority synthesis is explicit fan-out with citations. A shared slug or
semantic match never silently joins private and public corpora. Source: frozen
GBrain MEMORY_VERBS v1 and brains/sources topology at 75fae742, reviewed
2026-08-11
Runtime Memory Adapter Contract
Hindsight's current retain/recall/reflect system makes the portable job surface concrete, while also showing why runtime memory needs stronger boundaries than an ordinary vector-store adapter.
The adapter must preserve these distinctions:
| Boundary | Requirement |
|---|---|
| Source versus derived memory | Immutable source revisions stay outside the memory provider. Facts, observations, and mental models link back to provenance and remain replayable derivations. |
| Search versus synthesis | Recall is the default bounded evidence path. Reflect is an explicit, slower, inference-bearing operation with citations, model/cost receipt, and a degraded/unavailable state. |
| Bank versus tag scope | Use banks for hard user/authority/repository isolation. Shared banks require mandatory provenance tags, strict matching, and cross-scope leakage tests. |
| Local versus cloud | Record API/database, extraction model, embedding, reranker, and reflect provider independently. Loopback transport does not prove zero egress. |
| Export versus recovery | Document or whole-bank export is only the first half. Restore into a fresh target, rebuild derived indexes, rerun frozen queries, verify provenance/scope, and prove provider loss does not erase canonical evidence. |
| Runtime versus deployment truth | A selected provider is not active until pinned install, reversible wiring, doctor/readiness, smoke, leakage, usefulness, and recovery receipts exist. |
Hindsight v0.9.0 supports four-way recall (semantic, BM25, graph, temporal),
derived observations, standing mental models, agentic reflect, strict bank
boundaries, multi-harness per-repository routing, and document/whole-bank
transfer. Those features make it a strong adapter candidate, not an exemption
from the contract. Source: vectorize-io/hindsight@d7c33fde, reviewed
2026-08-11
Two Tiers
| Tier | Medium | Properties |
|---|---|---|
| Working memory | context window | fast, volatile, expensive, attention-limited |
| Long-term memory | wiki / DB / vector / graph / logs | durable, searchable, needs consolidation |
The architecture's job is not to remember everything. It is to move the right thing to the right tier at the right time.
Agent Company OS names the runtime version of this as a memory skill overlay: inject the relevant facts, examples, runbooks, and constraints for a worker's department without giving it the whole workspace. That is the practical boundary between useful long-term memory and context stuffing. Source: X/@DerekNee image, 2026-06-25
Storage Backends
- Document/wiki store: inspectable compiled truth and timelines.
- Vector store: fuzzy semantic recall.
- Keyword index: exact identifiers, commands, names, and rare terms.
- Graph layer: entities and relationships.
- Key-value store: exact settings, preferences, and state.
- Append log: replayable episodic history.
Production systems are hybrid. Kevin's wiki intentionally biases toward inspectable documents plus qmd retrieval, because the agent and Kevin both need to read the memory.
Latent Briefing adds a lower-level handoff path for multi-agent systems: instead of passing the orchestrator's whole reasoning trace as text, compact the relevant KV-cache state for the worker's task. That is not a replacement for durable memory; it is a runtime transfer optimization for "what the next agent needs right now." Use it as an inference/memory-system pattern, not as a wiki-storage pattern. Source: X/@RampLabs and technical source review, 2026-07-04
Write Path
Important invariants:
- Do not append duplicate facts when an existing page should be updated.
- Preserve evidence in timelines.
- Use aliases/frontmatter to resolve identity.
- Route recurring procedures to skills or automations.
- Rebuild indexes after broad writes.
Signal-driven maintenance loop
Memory is a governed feedback loop, not a passive cache:
- observe a complete source or execution trace and detect salient signals;
- create compact, provenance-bearing signals rather than another parallel page;
- bind them to stable objects and explicit graph relations;
- retrieve the smallest sufficient working set for a task;
- plan and act with that context;
- evaluate the resulting state or artifact; and
- update, supersede, merge, or prune derived memory from the outcome signal.
Kevin's deletion boundary is stricter than an ordinary agent cache. Immutable source revisions, attachment bytes, content hashes, and review receipts are not deleted because a compiled claim became stale or low-value. Maintenance may supersede a current claim, merge duplicate objects, retire a route, remove a derived projection, or drop an item from active retrieval while retaining the evidence needed to reconstruct and audit the decision. That preserves both failure recovery and the user's rule that useful saved signals do not disappear. Source: Self-Improvements in Modern Agentic Systems: A Survey, section 6.2; reviewed 2026-08-11
Module Decomposition
The Are We Ready For An Agent-Native Memory System? paper provides a system-level decomposition that should be used when designing or auditing memory products:
| Module | Design decision | Kevin proof |
|---|---|---|
| Representation and storage | Text, vectors, graph, temporal stream, hierarchy, or hybrid | source/artifact digests, object identity, recoverable raw evidence |
| Extraction | What gets written, when, and with what evidence | complete-source expansion, atomic signal coverage, provenance |
| Retrieval and routing | Which memory is selected for a task and how it enters context | recall at several budgets, evidence completion, distance/freshness tests |
| Maintenance | How memories are updated, consolidated, forgotten, or reorganized | digest-bound writeback, temporal correction, local replay and cost |
Do not evaluate "memory" as one feature. Evaluate the module that is likely to fail. In Kevin's wiki, extraction and maintenance are the risky modules; storage is intentionally boring markdown, and retrieval is qmd plus graph navigation. Preserve source context before abstraction, then filter at retrieval time. The paper's ablations show that summaries and deeper hierarchy cannot restore evidence removed during extraction. Source: arXiv 2606.24775v1, Tables 3–5, visually verified 2026-08-10
Read Path
- Search catalog/qmd.
- Read likely pages.
- Follow related/backlinks if the task crosses domains.
- Filter for current, source-backed claims.
- Inject only the useful slice into the active context.
This read path is Context Engineering in action: select before stuffing.
Every injection declares a mode: bootstrap for a small always-visible policy/profile set, pull for task-triggered search, or push for proactive selection. Bootstrap has a hard token and freshness ceiling. Pull is measured for know-to-ask misses. Push is measured for precision, recall, false fires, latency, decision delta, and injected tokens. Expensive synthesis or a post-turn reviewer is an explicit loop, never an invisible default on every synchronous turn. [Sources: MemGuide; PRIME, reviewed 2026-08-11]
Conflict Handling
Memory systems rot when they cannot represent changed facts. The wiki pattern handles this by separating compiled truth from timelines:
- Update compiled truth to the current best synthesis.
- Append a timeline entry for the new evidence.
- If an older claim was wrong, add a correction entry rather than silently erasing history.
- Merge duplicate entities instead of maintaining parallel pages.
Conflict handling also needs an explicit no-answer state. When the controlling source is missing, two current sources disagree, authorization removes needed evidence, or the retrieval budget cannot assemble sufficient support, the system should return the uncertainty and missing edge instead of converting a plausible guess into memory.
Evaluation
| Axis | Question | Representative metric |
|---|---|---|
| Write fidelity | Did capture preserve the full source and promote the right unit? | source coverage, unsupported-signal rate, raw recovery |
| Identity/consolidation | Did revisions and aliases bind to the right object? | duplicate/conflict rate, correction accuracy |
| Retrieval | Can the system assemble all controlling evidence? | recall@k, evidence-set recall, distance/freshness bins, token budget |
| Usefulness | Did memory improve the action rather than decorate the prompt? | task success, state correctness, decision delta |
| Temporal stability | Does memory remain correct after updates and long histories? | stale-fact rate, long-horizon degradation |
| Authority/abstention | Did scope apply before retrieval, and can the system decline? | leakage tests, conflict/no-answer accuracy |
| Cost/recovery | Is the value worth construction, query, and maintenance cost? | latency, token/API cost, update fan-out, replay success |
Evaluation has two non-substitutable layers. Conformance checks operation shape, error behavior, privacy, source isolation, budgets, idempotency, and write→read→expire round-trips. Usefulness checks whether memory surfaced when it should, stayed quiet when it should, volunteered the right context, preserved provenance, survived across harnesses, and improved the task without unbounded context injection. Passing conformance does not establish ranking quality; a good retrieval score does not prove privacy or writeback behavior.
The minimum cross-harness matrix uses the same synthetic fixture and frozen
budget across Codex, Claude Code, Hermes/OpenClaw, and any future adapter. It
reports know-to-ask failure and false-fire rates, push precision/recall,
writeback fidelity, provenance accuracy, continuity, source-isolation
violations, and injected tokens. Private scenarios must be invented rather
than name-swapped from real life, and held-out cases remain sealed from the
adapter. Source: GBrain BrainBench corpus, ledger, and baseline at 75fae742,
reviewed 2026-08-11
Memory benchmark comparison also freezes the memory release, corpus/split, answer model, extraction and preprocessing, query mode, recall/reflect budget, latency, cost, and scorer. Report retrieval evidence separately from final answer or task success. A project-published LongMemEval number cannot be ranked against another result when one side changes the answer model, pre-extracts the history, uses an agentic search loop, or runs at a different scale.
Freeze variant selection and judge policy too. Supermemory's published 98.6% eight-variant exercise is explicitly a parody: accepting any successful prompt variant demonstrates how configuration can manufacture a headline score. AMA- Bench adds real agent-environment trajectories and causal/objective retrieval; MemoryAgentBench separates accurate retrieval, test-time learning, long-range understanding, and conflict resolution. These are complementary lanes, not one leaderboard. [Sources: Supermemory; AMA-Bench; MemoryAgentBench, reviewed 2026-08-11]
The suite needs several distinct lanes: original LongMemEval for conversational extraction, temporal reasoning, knowledge update, and abstention; LongMemEval-V2-style cases for static/dynamic environment state, workflows, gotchas, and premise awareness; adversarial bank/tag leakage; recall versus reflect cost; no-memory and raw-context controls; and export-to-fresh-target replay. Hindsight's strong self-published results and retrieval-centered counterresults are both useful only after the protocol is attached. [Sources: Hindsight, LongMemEval-V2, and Storage Is Not Memory, reviewed 2026-08-11]
The paper's cost result is directionally useful: localized maintenance can beat global reorganization. That supports Kevin's owner-by-owner compiled-truth updates and targeted merge protocol over periodic whole-wiki rewrites. It also makes broad re-embedding or graph rebuilds derived maintenance jobs, not the semantic write path itself. Source: arXiv 2606.24775v1, Figure 11, 2026-06-23
Timeline
-
2026-08-12 | Implemented the first portable authenticated memory-write slice: token-derived authority, explicit source scope, owner- and provenance-bound proposed claims, hash-chained private storage, idempotent replay, scoped recall, evidence-preserving retirement, stable errors, and a live HTTP round-trip. Installed the same contract in the persistent Daily Brief LaunchAgent, added a portable TypeScript client, server-owned source allowlists and public-write policy, fsync durability, automatic proposal-only Daily Brief cards, and a repeatable conformance command. The installed write/replay/deny/retire/restart matrix passes while canonical wiki writeback remains behind source review. Source:
@kevin-wiki/brain-runtime;briefd;/api/memory; 24 runtime tests, 24 control-plane tests, and installed LaunchAgent conformance -
2026-08-11 | Added the optional bitemporal signal envelope, durable pre-compaction checkpoint, bootstrap/pull/push injection policy, and one-canonical-writer rule. Corrected unsupported product claims and added variant/judge sensitivity plus AMA-Bench and MemoryAgentBench evaluation lanes. Source: X/@michael_chomsky and primary-source audit, 2026-08-11
-
2026-08-11 | Added the runtime-memory adapter contract from the complete Hindsight v0.9.0 replay: immutable source versus derived facts/observations/mental models, recall versus reflect cost, hard bank and strict-tag scope, local-provider declaration, fresh-target export/replay, and selected-versus-active deployment state. Extended evaluation with protocol-matched LongMemEval, LongMemEval-V2 workflow/gotcha, leakage, cost, degraded-mode, and recovery lanes. Source: X/@itsharmanjot;
vectorize-io/hindsight@d7c33fde; LongMemEval/Hindsight/LME-V2/retrieval-centered papers -
2026-08-11 | Added a portable five-job agent-memory surface, two-axis authority/source scoping, the immutable-evidence boundary for remember/forget, and separate conformance versus cross-harness usefulness evaluations. Kept the authenticated writer/expiry interface explicitly unbuilt. Source: GBrain
MEMORY_VERBS v1, brains/sources topology, and BrainBench at75fae742 -
2026-08-11 | Mapped the paper's observe → compact → organize → retrieve → act → evaluate → update/delete memory cycle onto the evidence/object/proof architecture, with an explicit rule that pruning derived memory never erases immutable source revisions or decision receipts. Source: X/@omarsar0; arXiv:2607.13104v1
-
2026-08-10 | Added the lossless-evidence-to-finite-context layer model, module-specific proofs, explicit no-answer behavior, evidence-set retrieval metrics, and localized-maintenance cost rule from the complete paper and correlated infinite-context article. [Sources: arXiv 2606.24775v1; X/@daleverett, 2026-07-14]
-
2026-07-04 | Added Latent Briefing as the representation-level context handoff pattern for multi-agent memory: compact task-relevant KV-cache state instead of re-sending the full reasoning trace as text. Source: X/@RampLabs, 2026-04-10
-
2026-06-25 | Added the four-module data-management decomposition and evaluation axes from "Are We Ready For An Agent-Native Memory System?" Source: arXiv 2606.24775v1
-
2026-06-25 | Added Agent Company OS's memory skill overlay as the runtime pattern for injecting scoped memory into workers without broad context stuffing. Source: X/@DerekNee, 2026-06-25
-
2026-06-18 | Expanded with wiki-specific backend mapping, write/read paths, conflict handling, and evaluation axes. Source: User request, 2026-06-18
-
2026-06-02 | Page created. Captured working-vs-long-term tiering, vector/graph/KV/log backend mix, extract-consolidate-dedup write path, retrieve-rank-page read path, and durability implications. Source: compiled from Mem0, MemGPT/Letta, CoALA