Agent Self-Improvement Eval Library

Kevin's agent eval library turns a run into a receipt, scores the receipt against durable expectations, and evolves the evaluator when the agent either fails or saturates the current suite.

File-Native Structure

The local implementation follows the Vercel Eve lesson that agents should be visible as files:

evals/agent-self-improvement/
  suite.json                       # benchmark cases + controlled evaluator epochs
  examples/passing-receipt.json    # fixture receipt
scripts/run-agent-self-evals.ts    # scorer
scripts/lib/agent-self-eval-contract.ts
scripts/lib/__tests__/agent-self-eval-contract.test.ts
skills/productivity/agent-eval-library/SKILL.md
skills/productivity/loopy/SKILL.md

The runner is intentionally deterministic. It does not ask another model to decide whether the agent was good; it checks whether the receipt contains required evidence such as qmd search, skill reads, source verification, image inspection, routed page promotion, index rebuild, qmd embedding, routing doctor, and log writeback.

Why Receipts

An agent trajectory is too large to grade by memory. A receipt compresses the important path facts into a reviewable object:

  • events,
  • commands and exit codes,
  • files touched,
  • promoted pages,
  • notes about evaluator hardening.

That gives the evaluator a stable input. It also makes missing proof visible: if the agent claims it searched the wiki or inspected an image, the receipt must say so.

How To Run

npm run agent-self-eval
npm run agent-self-eval:test
npm run agent-self-eval -- --receipt evals/agent-self-improvement/examples/passing-receipt.json

The suite currently checks seven behaviors:

  1. brain-first search before external claims,
  2. skill-first routing before repeated-work codification,
  3. bookmark artifact promotion into durable pages,
  4. runnable eval proof,
  5. a bounded falsifiable quality loop,
  6. controlled evaluator epochs and checkpoint receipts, and
  7. graph closeout with index/qmd/routing/log evidence.

Evaluator Evolution

The suite implements controlled utility evolution from Red Queen Gödel Machine, not continuous judge mutation. Each active epoch freezes the evaluator ID/revision, artifact protocol, scoring rule, anchor revision, budget, and split. Failures and saturation may create challengers—a new deterministic check, negative fixture, atomic BINEVAL question, or cost/minimality constraint—but they cannot edit the active judge.

At a checkpoint, an acceptor separate from the proposer compares incumbent and challenger on the same evaluator-independent held-out anchor. A conservative lower bound governs promotion and ties keep the incumbent. Replacement preserves anchor evidence, invalidates records scored by the displaced evaluator, and forces the archive to re-rank under the new revision. Unanchored scores and winners are epoch-local; only fixed-anchor outcomes are comparable across revisions.

These are executable suite invariants. Five negative tests reject mid-epoch mutation, stale dependent scores, invalid cross-epoch comparison, and missing limitation boundaries. The full example receipt passes 23/23 assertions. This is the local answer to reward hacking without giving a moving evaluator permission to redefine success after seeing every output. Source: arXiv:2606.26294v2, Methods, Experimental Design, and Limitations; CaMLSys explainer, 2026-07-19; implemented 2026-08-11

Governed Critic And Trajectory Contract

Self-improvement is not a single high score. For each candidate scaffold revision, the eval receipt must retain:

  • baseline and every accepted or rejected iteration, not only the peak;
  • the fixed attempt/compute/API/tool/human budget;
  • held-out transfer checks and regressions on previously passing cases;
  • deterministic metrics before parameterized judges where possible;
  • judge identity, rubric version, judging budget, and repeated-run reliability when a model critic is necessary; and
  • the exact component changed so gains can be attributed rather than credited to the entire harness.

The update generator and acceptor are distinct roles. The generator may propose new tests, but it cannot both weaken the suite and approve the patch optimized against that suite. Red Queen evolution is monotone by default and checkpointed: queue a negative fixture, narrower assertion, or proven threshold for anchor-backed comparison while the current epoch stays frozen. A deletion or relaxation needs an independently reviewed reason and audit trail. This makes the evaluator governed infrastructure rather than a self-confirming prompt.

Kevin's current program is scaffold improvement: prompts/instructions, memory, tool routing/refinement/creation, and full control logic can change while model weights remain fixed. The companion survey repository is retained as a watched research corpus, not installed as runtime code. Its taxonomy should classify future eval cases and help find primary papers; a bibliography entry alone is not proof of a method's result. Source: survey, sections 3, 6, 8, and 9; companion repository at 57a1d89e5bafcd65db7feb51e809422700aeb48a; reviewed 2026-08-11

Eval Source Map

The Awesome Evals thread and repo sharpen the local suite's design rule: every check must name its evaluated object before choosing a scorer. Final text can use binary rubric questions, structured output can use deterministic assertions, agent work should prefer trajectory receipts and state diffs, and self-improvement needs evaluator-evolution cases that the current judge would miss. Source: X/@xdotli, 2026-06-24; Source: GitHub benchflow-ai/awesome-evals, accessed 2026-06-30

For Kevin's agent receipts, this means the suite should keep adding narrow checks for actual misses: missing wiki search, unread skill, skipped visual artifact, unstated version snapshot, weak source verification, missing log/index/qmd closeout, or claims that a broad doctor proves a narrower requirement it does not cover. Those are trajectory failures, not prose failures, so they belong in deterministic receipt checks before any LLM judge.

Cross-harness transcript and environment receipt

An eval adapter normalizes each harness into one canonical receipt without flattening failures: ordered messages, reasoning/result boundaries where the harness exposes them, tool name and arguments, tool result, loaded skill or documentation reference, file/state diff, failed turn, process exit, retry, terminal assertion, tokens, latency, and cost. A zero process exit cannot override a failed agent turn, and a polished final answer cannot override the wrong tool call or unchanged target state.

Each run also labels environment fidelity as emulated-api, local-lite, local-real, or hosted-real, with exact image/service revisions, network and credential policy, seeded state, ports, Docker/socket access, cleanup, and the checks actually executed. Track tool-called, tool-avoided, documentation use, and skill activation as diagnostic mechanisms; accept changes only on the terminal task/state scorer and held-out regressions. Preserve the first attempt and every retry separately. These rules adapt the strongest reusable pieces of Supabase Evals without importing its product-specific scenarios. Source: supabase/evals@5f8cfeb9, reviewed 2026-08-12

Skill-promotion receipts

Repeated run traces may propose a reusable skill, but the eval layer owns the promotion gate. A skill-promotion receipt must identify supporting and contrary runs, its trigger and exclusions, applicability boundary, deterministic or reviewable verifier, fallback path, and the held-out cases used to estimate reliability. Newly generated skills start probationary. Only task-specific proof moves them into active routing; failure narrows or revises the skill, and low reliability or inactivity removes it from active retrieval without deleting its evidence.

Reliability is a routing statistic, not causal credit. A successful terminal reward may co-occur with an irrelevant step, and reflection can be wrong. The suite therefore prefers state diffs, deterministic checks, counterexamples, and held-out runs over self-reported confidence. Source: MSCE paper, sections 4.3, 5.4, and Limitations; reviewed 2026-08-10

Controlled skill-optimization lane

Skill evolution keeps the evaluator fixed and changes one external skill artifact. The controlled-skill-optimization/v1 contract freezes the target model, execution harness, evaluator, split revisions, and incumbent digest; uses training traces only to create bounded add/delete/replace patches; selects only on a disjoint held-out split with strict improvement and tie rejection; and reserves the sealed final test for reporting. Holdout leakage causes abstention, not an “unverified” promotion.

The lane grades frontmatter routing, activated-body task behavior, and normal end-to-end activation separately. It reports per-skill effect sizes, regressions, false triggers, and negative transfer rather than relying on a corpus average. Rejected patches and score deltas remain negative evidence; protected slow policy cannot be changed by ordinary fast edits. Canonical source changes only after a reviewed provenance-bearing diff with rollback. Open-ended work stays held without a reliable verifier or review evidence.

Five negative tests now reject leaked holdouts, tied lateral drift, description/body conflation, aggregate-only promotion, unbounded full rewrites, mutable slow state, and auto-promotion of unverifiable open-ended work. The contract deliberately does not install SkillOpt-Sleep or grant a provider budget; it adapts the paper and router-benchmark evidence to the existing skill-creator and eval stack. [Sources: SkillOpt v2; Context Engineering Agent Skills v2.3.0; implemented 2026-08-11]

Guardrail-policy optimization lane

A guardrail policy needs a two-sided acceptance rule. The candidate objective is not merely “reduce attacks”; it is “reduce attack success relative to the unguarded target while preserving legitimate behavior.” Each receipt must bind:

  • one mutable policy digest and an unchanged harness, target, judge, judge prompt, suite, split, and wall-clock/API budget;
  • attack success with no policy, attack success with the incumbent, and attack success with the candidate;
  • benign-pass or task-success rate for the same three arms, including a declared maximum permitted degradation;
  • repeats or variance whenever either model path is stochastic;
  • the exact keep/reject decision, rejected candidate, restored incumbent, and rollback digest; and
  • a new lineage whenever the suite, judge, harness, or task boundary changes.

This prevents the classic all-refusal reward hack and separates native model safety from the marginal value of the policy. It also limits causal claims: a single-turn text suite cannot prove tool, file, multi-step, or production-agent safety. Santander's autoguardrails v0.1.0 is a compact executable reference for this shape—140 fixed cases, one mutable policy, an unguarded baseline, append-only decisions, and automatic restore after rejection—but its shipped perfect score is against a deterministic local stub. It is harness proof, not a production benchmark. Twenty-four stdlib-compatible upstream tests passed in the pinned replay; two pytest-only detector files were not executed. Source: SantanderAI/autoguardrails@1ca0c9b, reviewed 2026-08-11

Case-adaptation eval receipts

An adaptive harness changes the evaluated object from one global scaffold to a pair: default harness + case delta. Its receipt must therefore bind every outcome to the default digest, the visible pre-run case profile, the retrieved experience IDs, and the exact changes across context, tools, generation, orchestration, memory, and output processing. The evaluator rejects a claimed feedback-free run when the adapter could observe the current label, verifier result, or post-outcome state.

The minimum paired evaluation compares the unchanged default against the case-conditioned arm under the same model, tools, runtime, budget, split, and verifier. Then Harness Ablation Design removes each control surface and the per-case, global-pattern, and adaptation layers separately. Receipts retain negative transfer, regressions, uncertainty, and cache-aware cost telemetry; high raw-token retrieval is not called cheap merely because one experiment cached most of it.

The MemoHarness reference code is retained as a watched research artifact, not installed. It is a v0.1.0 Harbor/Daytona implementation with no declared license, and the reviewed repository's test surface is too narrow to establish runtime fitness. Its paper supplies a valuable experiment shape while candidly leaving statistical robustness and component attribution open. Source: MemoHarness paper, especially sections 3–4 and Appendix A–C; repository at e7da7728b6a8020f7e13c890d50df32ad6f2710c, reviewed 2026-08-11

Cross-harness memory conformance lane

Memory requires a dedicated suite rather than another assertion inside the generic agent-run receipt. The lane has two gates:

  1. Protocol conformance: required fields, enum/error behavior, mandatory provenance, visibility and source isolation, budget arithmetic, miss behavior, idempotent expiry, and a deterministic write→read→expire round-trip.
  2. Behavioral usefulness: know-to-ask failure and false-fire rates, volunteered-context precision/recall, provenance-bearing writeback, cross-harness continuity, source-isolation violations, and injected-token volume.

Fixtures must expose only adapter-visible conversation data; sealed gold stays separate. At least one held-out partition is excluded from the ordinary CI gate. Synthetic cases use invented people, companies, relationships, amounts, and timelines rather than anonymized real scenarios. A harness comparison freezes the memory backend, corpus, task, budget, model when possible, adapter version, and scoring code so the result measures the seam instead of a moving stack.

Kevin-Wiki now passes the first gate for its declared persistent single-writer boundary. The installed Daily Brief LaunchAgent and MemoryHttpClient prove server-derived authority, mandatory provenance and owner binding, exact source scope, public-write denial, idempotent remember/replay, bounded recall, retirement, restart durability, retained history, and proposal-only review projection. This is protocol conformance, not a fake usefulness result.

The second gate remains open: volunteered-context precision/recall, know-to-ask and false-fire behavior, injected-token cost, and continuity across executed Codex, Claude, Cursor, and Hermes task fixtures still need a sealed corpus and frozen evaluator. That behavioral suite cannot graduate with any source-isolation violation. Source: GBrain BrainBench and MEMORY_VERBS v1 at 75fae742; scripts/memory-conformance.ts; installed replay, 2026-08-12

Memory benchmark methodology gate

A memory score is not comparable until the receipt freezes the complete evaluated stack:

Field Required value
Memory implementation product/repository and exact release or commit
Corpus dataset, split, scale, history construction, and contamination policy
Answer model model/provider/version and decoding parameters
Ingestion raw versus summarized input, extraction model, prompt, and preprocessing
Query path recall, reflect/agentic search, full context, or another named mode
Budget retrieved items/tokens, context, tool iterations, time, and retries
Scorer exact evaluator, judge model/version, normalization, and abstention policy
Operations hardware/service topology, cache state, concurrency, and failure handling
Cost ingest, maintenance, query, and export/recovery cost reported separately
Provenance evidence returned, scope applied, and source/revision identity

The receipt reports retrieval/evidence quality separately from final answer or task success. It also records self-published versus independently reproduced results. A higher number under a different answer model, preprocessing pass, agentic loop, scale, or scorer is a new experiment—not a leaderboard update.

The minimum test family is broader than conversational question answering:

  1. original LongMemEval-style extraction, multi-session reasoning, temporal reasoning, knowledge update, and abstention;
  2. LongMemEval-V2-style static/dynamic environment state, workflow knowledge, gotchas, and premise awareness;
  3. no-memory, raw-context, recall-only, and reflect controls under one budget;
  4. stale/corrected fact, contradiction, and temporal-validity cases;
  5. strict bank/tag source-isolation and malicious cross-scope queries;
  6. know-to-ask versus volunteered-context usefulness and intrusion accounting;
  7. unavailable model/provider and partial-index degraded modes; and
  8. export to an empty target followed by provenance, retrieval, and behavior replay.

This gate keeps Hindsight's strong current self-published results useful without turning them into a universal winner claim. The Hindsight paper, later project benchmarks, LongMemEval-V2, and retrieval-centered competing results usefully disagree because they expose how much the answer depends on the protocol. [Sources: LongMemEval, Hindsight, LongMemEval-V2, and Storage Is Not Memory, reviewed 2026-08-11]

Concept Position

Field Value
Concept family Brain, memory, and retrieval
Concept owned Kevin's agent eval library turns a run into a receipt, scores the receipt against durable expectations, and evolves the evaluator when the...
Category map Concept System Map

Timeline

  • 2026-08-11 | Added a guardrail-policy optimization lane with an unguarded baseline, paired attack/benign objectives, fixed lineage, rejected-candidate retention, rollback, and explicit single-turn limits. Source: SantanderAI autoguardrails@1ca0c9b; exact-source replay

  • 2026-08-11 | Added the machine-checkable controlled-skill-optimization lane: frozen run identity, train/selection/final-test separation, leak abstention, bounded patches, protected slow policy, rejected-edit evidence, separate description/body/end-to-end measurement, per-skill effects, strict improvement, open-ended holds, reviewed promotion, and five negative tests. Source: X 2059113412278227328; SkillOpt arXiv v2 and source; Context Engineering Agent Skills v2.3.0

  • 2026-08-11 | Replaced slogan-level Red Queen hardening with a machine-checkable controlled-utility-evolution contract: frozen evaluator epochs, anchor-revision identity, independent challenger acceptance, tie-to-incumbent, selective invalidation, fixed-anchor-only cross-epoch comparison, five negative invariant tests, and a 23/23 example receipt. Source: X 2071285506630160761; arXiv:2606.26294v2; CaMLSys explainer

  • 2026-08-11 | Added a protocol-complete memory benchmark gate from the Hindsight replay: exact memory/corpus/model/ingestion/query/budget/scorer/operations/cost/provenance identity, retrieval-versus-task separation, self-published labels, workflow/gotcha and privacy lanes, degraded modes, and fresh-target export/replay. This prevents unmatched LongMemEval or BEAM numbers from silently becoming a provider default. Source: vectorize-io/hindsight@d7c33fde; LongMemEval, Hindsight, LongMemEval-V2, and Storage Is Not Memory

  • 2026-08-11 | Added the cross-harness memory lane: separate protocol conformance from behavioral usefulness; require sealed gold, holdouts, invented scenarios, fixed comparison conditions, source isolation, intrusion accounting, and honest unbuilt status for Kevin's writer/expiry/adapters. Source: GBrain BrainBench and MEMORY_VERBS v1 at 75fae742

  • 2026-08-11 | Added case-adaptation eval receipts binding outcomes to the default digest, pre-run evidence and six-surface delta; required same-condition paired controls, leakage rejection, component ablations, negative transfer, uncertainty, and cache-aware cost; retained MemoHarness as watched research code rather than an installed dependency. Source: X/@omarsar0; arXiv:2607.14159v1; HowieHwong/MemoHarness

  • 2026-08-11 | Added the governed-critic and trajectory contract: fixed budgets, held-out and regression checks, full iteration history, evaluator identity/reliability, component attribution, proposer/acceptor separation, and monotone evaluator changes. Preserved the survey PDF, source image, project page, and MIT companion repository snapshot. Source: X/@omarsar0; arXiv:2607.13104v1; selfimproving-agent repository

  • 2026-08-10 | Added probationary skill-promotion receipts with evidence anchors, applicability boundaries, verifiers, fallbacks, held-out cases, and reliability lifecycle checks; retained the paper's warning that heuristic value is not causal credit. Source: X/@dair_ai; arXiv:2607.16621

  • 2026-07-01 | Concepts category refresh added this page to the Brain, memory, and retrieval family, linked it to Concept System Map, and kept it standalone because it owns this reusable mental model: Kevin's agent eval library turns a run into a receipt, scores the receipt against durable expectations, and evolves the evaluator when the... Source: User request, 2026-07-01

  • 2026-06-30 | Folded the Xiangyi Li / BenchFlow eval source map into the local suite model: choose scorers by evaluated object, prefer deterministic/state/trajectory checks before LLM judges, and reserve Red Queen evaluator evolution for self-improving agents. Source: X/@xdotli, 2026-06-24; GitHub benchflow-ai/awesome-evals, 2026-06-30

  • 2026-06-29 | Created the first local self-improvement suite and runner after Kevin asked Codex to build an eval library to benchmark itself and self-improve. Source: User request, 2026-06-29