Agent Memory Patterns
Memory turns repeated agent work into compounding system behavior. The pattern is extract, consolidate, store, retrieve, and act.
Short-term memory is the finite working set in the live context window. Long-term memory is everything durable enough to survive a new session: source artifacts, wiki pages, logs, skills, rules, project docs, and structured stores. Addressing a large store is not the same as placing it all in context. The important design questions are: what kind of memory is this, who owns its write path, which evidence may enter the current working set, and how can the agent know when the store does not contain an answer?
Three Memory Types
| Type | Meaning | Kevin-stack examples |
|---|---|---|
| Semantic | durable facts and synthesized knowledge | wiki compiled truth, project pages, tool pages |
| Episodic | records of specific events | timelines, wiki/log.md, activity logs, transcripts |
| Procedural | how to do work | skills, workflows, rules, doctors, automation prompts |
Confusing these types causes rot. A timeline entry should not overwrite compiled truth by itself. A one-off chat note should not become a rule. A useful procedure should not sit as trivia on a project page.
Archive Is Not Canonical Memory
Kevin's system separates retention from promotion:
- The evidence archive preserves every captured source revision and its media, links, provenance, visibility, and content digest.
- Atomic signals preserve potentially useful claims, preferences, questions, tasks, and counterexamples without pretending they are all true.
- Stable operating objects hold the smaller set of current instructions, workflows, skills, design rules, knowledge, projects, and decisions that survived interrogation.
- The active context receives only the bounded slice needed for the task.
This resolves the apparent conflict between “store every message” and “do not treat every message as permanent truth.” Capture should be lossless. Canonical promotion should be selective, source-bound, and reversible.
The @trq212 implementation-notes.html prompt is a live hybrid of episodic and procedural memory: while the agent implements a spec, it keeps a running note of ambiguous decisions, necessary deviations, tradeoffs, and anything the human should know. That gives the model permission to make local decisions without hiding them, and it gives the next reviewer a durable diff between the requested spec and the system that actually emerged. Source: X/@trq212, 2026-05-18
Production Memory Pipeline
The hard part is consolidation. Appending everything is easy and bad. Good memory updates the existing page, merges aliases, preserves evidence in timelines, and changes the compiled truth only when the evidence justifies it.
Wiki Mapping
Kevin's wiki is a deliberately human-readable memory system:
- Compiled truth = semantic memory.
- Timeline = episodic memory.
- Skills and workflows = procedural memory.
- Frontmatter and aliases = identity resolution.
- Backlinks and related fields = graph memory.
- qmd = retrieval.
- doctor and routing evals = memory integrity checks.
This design trades automation glamor for inspectability. A vector store can retrieve facts, but a wiki page can explain why the fact matters and how it changed over time.
Runtime Memory Boundary
Kevin's agent stack now separates memory by authority:
- qmd/Kevin-Wiki owns durable compiled truth.
- Hindsight owns Hermes runtime recall and graph-shaped reflection.
- Honcho owns profile/user/peer modeling.
- Graphify owns repo topology and affected-path reasoning.
- Agent docs own project-local working facts and procedures.
The pattern: retrieve from the narrowest memory lane, verify at the source when the answer matters, then write back any reusable lesson to the durable layer. Runtime memory can suggest; the wiki, skill tree, automation files, config, and project docs decide.
Write Discipline
Agents should write memory when:
- Kevin asks to remember, ingest, add, compile, or update.
- A repeated workflow appears.
- A decision, postmortem, or contradiction emerges.
- A project/tool/person page is stale after new evidence.
- A skill or rule would prevent future recurrence.
Agents should not write memory for transient task state, guesses, or unverified external claims.
Retrieval Discipline
Before answering from memory:
- Search qmd or
_index.md. - Read the likely page plus backlinks/related pages.
- Follow only high-signal links.
- If current facts might have changed, verify against primary sources.
- File the reusable synthesis back to the wiki.
Why Structured Memory Beats Giant Context
Stuffing all history into a large context is expensive, slow, and noisy. Structured memory lets the agent select the right unit at the right time and preserve contradictions instead of blending them into soup. The goal is not perfect recall. The goal is useful, current, inspectable memory.
Memory As A Data System
The 2026 paper Are We Ready For An Agent-Native Memory System? sharpens this page's thesis: agent memory should be evaluated as a data system, not a magic context appendage. Its module split is a useful checklist for Kevin's wiki and any product memory layer:
| Module | Wiki equivalent |
|---|---|
| Representation and storage | markdown pages, frontmatter, aliases, qmd index, backlinks |
| Extraction | ingest scripts, bookmark enrichment, transcript mining, source compilation |
| Retrieval and routing | qmd search, _index.md, active-stack/router pages, related links |
| Maintenance | doctor checks, timeline/compiled-truth separation, stale-source checks, dedup/merge protocol |
The paper's most useful operational claim is that architecture should match the workload bottleneck. A memory system for exact tool names needs keyword and alias discipline; one for fuzzy project recollection needs semantic retrieval; one for evolving facts needs update correctness and maintenance policy. Source: arXiv 2606.24775v1, 2026-06-23
Its component ablations add four stronger defaults:
- Preserve recoverable source evidence before summarizing it.
- Filter late: capture broad context at write time, select narrowly at read time, and never confuse the two stages.
- Treat early localization and complete evidence assembly as separate retrieval targets.
- Consolidate conservatively and locally; broad rewriting can lose details and become the dominant cost.
Memory Behavior Contract
A durable memory system should prove all of the following:
| Behavior | Required proof |
|---|---|
| Write fidelity | source revision and raw artifact remain recoverable after extraction |
| Identity | a correction binds to the same entity/event instead of becoming an unrelated fact |
| Evidence assembly | all controlling support appears within a tested retrieval budget |
| Temporal correctness | the current valid state outranks stale mentions without deleting history |
| Conflict/no-answer | the agent exposes disagreement or abstains when evidence is insufficient |
| Authority | access scope is applied before retrieval, fusion, or graph expansion |
| Maintenance | updates are local, replayable, and digest-bound |
| Behavior change | retrieved memory measurably changes the next decision or action correctly |
Dale Everett's “infinite context” article reinforces the read-side rule: store, search, select, and inject the few pieces that matter, with dates, sources, approval state, access control, conflict tests, and a no-answer path. The phrase describes moving selective attention over durable evidence, not an infinite prompt and not a default recommendation to adopt pgGraph or Polygres. Source: X/@daleverett, 2026-07-14
Files Are Authority; Engines Are Derived
Michael Chomsky's critique of the “memory is markdown” school preserves a useful counterweight: portable files are an ownership decision, not a complete memory implementation. Identity, temporal validity, retrieval, access control, collaboration, compaction, and evaluation remain system responsibilities. Source: X/@michael_chomsky, 2026-04-12
The practical rule is to separate authority from projection:
- markdown, JSON, source artifacts, and review receipts can remain the inspectable, versioned authority;
- keyword, relational, vector, and graph stores may be rebuildable engines;
- an engine is not a second authority unless it owns a documented write class;
- deleting or rebuilding an index never deletes canonical evidence.
GBrain demonstrates the compatible hybrid: markdown is the system of record
while PGLite/Postgres and pgvector power derived search. That is not a failure
of file ownership; it is a reminder to make projection replay explicit.
Source: frozen GBrain repository at
75fae742d55ade4b29c1a574eb7b62bf5b053518, reviewed 2026-08-11
Three Context-Injection Modes
| Mode | Use | Main failure | Required proof |
|---|---|---|---|
| Bootstrap | Small identity, policy, and active-project memory needed every turn | context bloat and stale always-on rules | token ceiling, freshness, and removal test |
| Pull | Agent searches when it knows what it needs | “did not know to ask” | frozen-query recall and no-answer behavior |
| Push | System predicts useful memory before drift or omission | false fires, latency, and extra inference | push precision/recall, decision delta, and injected-token cost |
MemGuide's marginal slot-completion gain and PRIME's proactive memory evolution are research directions for push selection, not permission to add an inference round to every synchronous turn. A post-turn reviewer such as Saguaro is a feedback/verification loop that may detect missing context; it is not by itself a complete memory system. [Sources: MemGuide; PRIME; Mesa Saguaro, reviewed 2026-08-11]
Memory Is Still A Tradeoff System
Portable memory is the correct ownership posture, but the cheapest adequate mechanism should win for each workload. A file may beat a graph for a short, stable policy; a temporal graph may beat a flat note for changing relationships; bounded lexical search may beat inference for exact commands. The burden is behavioral proof, including cost—not architectural novelty.
Benchmark labels do not remove that burden. MemoryAgentBench evaluates accurate retrieval, test-time learning, long-range understanding, and conflict resolution; “selective forgetting” is a useful maintenance goal but is not the paper's fourth named ability. Supermemory's 98.6% eight-variant post is explicitly a parody and demonstrates scorer/configuration sensitivity, not a generally superior production system. [Sources: MemoryAgentBench; Supermemory, reviewed 2026-08-11]
Timeline
- 2026-08-11 | Re-reviewed the complete memory critique and primary sources; separated file authority from derived search engines, added bootstrap/pull/push injection modes, corrected MemoryAgentBench's fourth ability to conflict resolution, and treated the 98.6% Supermemory post as a benchmark-protocol warning rather than a product ranking. Source: X/@michael_chomsky and frozen audit, 2026-08-11
- 2026-08-10 | Reconciled lossless source retention with selective canonical promotion; added late filtering, complete-evidence retrieval, temporal/no-answer behavior, local maintenance, and the moving-working-set interpretation of “infinite context.” [Sources: arXiv 2606.24775v1; X/@daleverett, 2026-07-14]
- 2026-07-03 | Deep-reviewed the @trq212 self-thread and clarified implementation notes as live episodic/procedural memory during agent implementation. Source: X/@trq212, 2026-05-18
- 2026-07-04 | Added Michael Chomsky's critique of oversimplified agent memory: ownership can be markdown/git-backed, but production memory still needs tradeoff management, observability, collaboration semantics, update policy, and benchmarks. Source: X/@michael_chomsky, 2026-04-12/13
- 2026-06-25 | Added the agent-memory-as-data-system lens from "Are We Ready For An Agent-Native Memory System?" and linked the MemoryData benchmark suite. Source: arXiv 2606.24775v1
- 2026-06-18 | Expanded with Kevin's semantic/episodic/procedural wiki mapping, write discipline, retrieval discipline, and memory pipeline. Source: User request, 2026-06-18
- 2026-05-31 | Page created from Mem0's long-term-memory writeup and the CoALA memory taxonomy. Captured the short vs long-term split, triad, extract-consolidate-store-retrieve pipeline, and token/latency case for structured memory. Source: Mem0, https://mem0.ai/blog/long-term-memory-ai-agents, 2026-05-31
- 2026-05-18 | @trq212's "running implementation-notes.html" technique as a procedural/episodic memory store. Source: X/@trq212, 2026-05-18