MemoryData

Research harness for comparing heterogeneous memory methods under one launcher and artifact layout. Retained for evaluation; not installed.

MemoryData is the code artifact for Are We Ready For An Agent-Native Memory System?. At reviewed HEAD bdbe698, the repository exposes four benchmark families, 22 method presets, and a common main.py runner. The paper reports five workload families and 11 datasets; that experimental claim is broader than the four packaged repository families and should not be collapsed into the README tagline. Source: MemoryData README and arXiv 2606.24775v1, reviewed 2026-08-10

What It Can Compare

Family Primary stress
MemoryAgentBench accurate retrieval, conflict resolution, test-time learning
LoCoMo long-dialogue episodic, temporal, open-domain, and multi-hop QA
LongBench context-length robustness on the packaged proportional subset
MemBench recall, noise, knowledge update, high-level reasoning, multi-session recommendation

The presets span long-context, BM25/dense RAG, sequential memory, graph/tree memory, and hybrid approaches. Results, agent state, and logs are written under a stable artifact root, which makes the harness useful when a real project needs comparative memory evidence rather than a vendor demo.

Smallest Safe Evaluation Shape

python main.py \
  --agent_config config/reference_long_context_agent.yaml \
  --dataset_config benchmark/memoryagentbench/Accurate_Retrieval/config/EventQA/Eventqa_full.yaml \
  --max_test_queries_ablation 1 \
  --artifact_root /tmp/memorydata-smoke

Run a benchmark only in an isolated environment with pinned model, embedding, dataset, method, and repository revisions. Record endpoint/provider, prompts, model parameters, input dataset digest, output artifact digest, wall time, token/API cost, and failures. Compare at least one simple baseline against the candidate; a single complex method cannot establish improvement.

--retry_failed_queries resumes failed rows. --force deletes supported saved results and state before rebuilding, so it is not a harmless retry switch.

Adoption Boundary

MemoryData is not a production memory service and does not belong in the wiki runtime. The current repository requires Python 3.11, a large Python/ML dependency surface, external datasets, and model/embedding endpoints. Several presets assume additional method-specific runtimes or persistence.

At HEAD bdbe698f776d921ac791d1b07c0a7fc65a8bb4bb:

  • the root has no license file and GitHub reports no detected license;
  • there are four commits and no release tags;
  • no committed test suite or GitHub Actions workflow was found;
  • datasets are not bundled;
  • the default dependency manifests are broad rather than a hermetic lockfile.

Therefore the source was archived and inspected but not installed, imported, or executed. Reopen adoption only when a project has a concrete memory comparison question, upstream licensing is clarified, the minimal benchmark slice and datasets are pinned, and the smoke run produces a reproducible cost/quality receipt.

What To Measure

Do not reduce the run to answer F1. Record:

  • task outcome and executable state correctness where applicable;
  • evidence recall at several budgets and across evidence-distance bins;
  • stale-fact suppression after corrections;
  • long-history degradation;
  • write/index construction time, query latency, and cost;
  • provenance, abstention, and access-control behavior;
  • maintenance scope and recovery after interrupted or repeated runs.

Preserved Snapshot

The paper PDF, text extraction, visually inspected result pages, README, requirements, runner, and repository archive are preserved under .brain/artifacts/x/2069846777977880769/agent-memory-data-system/ with SHA-256 receipts. The source archive SHA-256 is b2fef5b0657115155ff518c2d29e8b71542f061faf68c7c8278c3a2764f76bea.


Timeline