Workflow Run Contract

A workflow run is complete only when a future agent can see what ran, what was promoted, what was deferred, what state changed, and which generated surfaces were refreshed.

This is the common completion bar for every page in wiki/workflows/. Individual workflow pages own domain-specific sources and decisions. This page owns the invariant that the run leaves proof.

Routing

Open this page when:

  • running any workflow manually
  • writing or updating a workflow page
  • converting a workflow into an automation
  • deciding whether a workflow result should update durable wiki pages
  • auditing whether a scheduled job actually completed

Completion Invariant

Every workflow run must make five decisions explicit:

Decision Required proof
What sources were read? Source list, command output, URLs, raw files, or repo paths.
What changed? Changed wiki pages, generated files, code files, or "No durable update needed."
What did not get promoted? Deferred items, blockers, low-signal findings, or no-op rationale.
What state advanced? state.json, raw cursor files, sync timestamps, scheduler status, or no state change.
What validation ran? Commands, checks, qmd queries, screenshots, test output, or explicit skipped-check rationale.

Standard Run Shape

  1. Read the workflow page and any canonical automation or skill it references.
  2. Confirm sources, credentials, repo path, and side-effect boundaries.
  3. Run the ordered steps, keeping raw/source material separate from compiled wiki truth.
  4. Write output under outputs/<YYYY-MM-DD>/<slug>/<agent>.md when the workflow is scheduled, long-running, or evidence-heavy.
  5. Promote only facts that pass the page's promotion criteria.
  6. Update state.json only for work that truly completed.
  7. Append wiki/log.md for durable or user-visible work.
  8. Follow Generated Surface Contract for indexes, qmd, context projections, registries, redirects, and stateful ledgers.

Immutable workflow-plan preflight

Before execution, resolve the job through brain/workflows.json and freeze the selected version. The plan records how it matched, who invoked it, its declared status and authority, input references, owner and progressive bindings, ordered pending stages, pending proof checks, and writeback targets. It carries a digest for the selected manifest, the complete registry, and the complete plan; the run ID derives from that plan digest.

Reject the plan before work when its digest no longer verifies, its manifest or registry changed, an agent selected a user-only workflow, a required binding is missing, or a stage/proof item was pre-marked complete. A prompt can provide inputs and case context but cannot replace the authority, reorder stages, delete proof, or broaden writeback without producing a newly reviewed plan.

npm run brain -- workflow route "<request>" [--agent]
npm run brain -- workflow plan <id-or-request> [input-ref...] [--agent]
npm run brain -- workflow begin <id-or-request> [input-ref...] [--agent]

Durable workflow-run ledger

workflow begin is the execution boundary. It verifies and freezes the same plan as workflow plan, then appends a run.started event to .brain/workflow-runs.jsonl. Every later transition is another immutable event; the ledger is not an editable status document. Events carry stable IDs, idempotency digests, prior-event hashes, and event hashes. Writes use an exclusive lock, append, and fsync; replay recomputes the chain and rejects a changed, reordered, duplicated, or detached event.

The run reducer enforces these rules:

  1. stages complete in manifest order and cannot be skipped;
  2. every proof attempt records the check, result, evidence, and time, so a failure followed by a pass remains visible;
  3. complete is illegal until every stage is complete and every required proof check has a latest passing attempt;
  4. completion records all five decisions from this contract: sources, changes, deferred work, state movement, and validation;
  5. failed, cancelled, expired, exhausted, stalled, and superseded outcomes remain honest terminal states with a reason and resume condition; and
  6. terminal runs are immutable. A retry begins a new run rather than rewriting history.

Use the CLI to append a checked event, reconstruct a receipt after a restart, list runs, or verify the entire chain:

npm run brain -- workflow record <event-json-path>
npm run brain -- workflow receipt <run-id>
npm run brain -- workflow runs
npm run brain -- workflow doctor

The receipt is derived state: selected manifest and plan digests, ordered stage history, proof-attempt history, terminal decisions, evidence, and a receipt digest. The JSONL ledger stays private because inputs and evidence may name local sources. Public interfaces may expose sanitized counts and declared workflow structure, but must not publish raw run evidence or treat a frozen plan as a completed run.

Release and self-update proof profile

Treat a release as a staged authority chain, not one successful command. A run may inspect or prepare without permission to publish; build without permission to sign; sign without permission to upload; publish one target without access to every target; and check for an update without permission to replace the installed executable.

Before an external release or updater mutation, freeze and prove:

  1. source repository, branch/tag/commit, clean-worktree policy, release version, changelog inputs, and the digest of the release configuration;
  2. the exact prepare preview: version changes, generated changelog, release branch/commit/tag plan, target registry/account/project/channel, and expected current remote state;
  3. every artifact's platform/architecture/runtime requirements, byte count and digest, dependency lock, SBOM/provenance, signing/notarization state, and the CI run that produced it;
  4. platform install/start/help/smoke tests plus compatibility refusal for an unsupported OS, architecture, libc, asset, signature, or runtime;
  5. the separately approved publication plan and its hash, approver and expiry; publication receipts must name each uploaded target, immutable version, remote digest, release URL, and partial-failure state;
  6. for delta delivery, exact base and target digests, authenticated manifest, patch-chain identity/order, per-hop and cumulative verification, size and decompression/resource caps, safe temporary destination and atomic replace, full-download fallback, interruption recovery, and rollback artifact; and
  7. post-release verification, revoke/yank limits, rollback command/artifact, credential cleanup, and the durable owner/state/log writeback.

Framework choice stays proportional. Ordinary project scripts remain first; Stricli is conditional on complex typed CLI routing, official generated clients on repeated service API work, Craft on real multi-target publication, Fossilize on chosen Node SEA distribution, and binpatch on measured binary-update cost. None of these tools expands authority. Enriched child links are discovery hypotheses until the primary repository is resolved. Source: X 2082947912498024751–2082947922648281278; pinned BYK/binpatch, BYK/fossilize, getsentry/sentry-api-schema, bloomberg/stricli, and getsentry/craft explorations, 2026-08-12

Proof-Driven Quality Loop

Quality-sensitive workflows add a bounded inner loop:

  1. freeze binary acceptance checks and a reproducible baseline;
  2. map coupled versus independent concerns before assigning work;
  3. change one highest-impact failure;
  4. run deterministic doctors/tests, then runtime or visual comparison;
  5. ask a fresh critic to find counterevidence and test whether the requested direction is itself wrong;
  6. keep only evidence-improving changes and convert recurring failures into gates;
  7. stop at 100% of declared checks, a proven hard ceiling/blocker, or two passes without meaningful improvement, recording the honest residual gap.

Parallel workers are appropriate for independent owners. Coupled rendering, state, data, or policy systems need a sequential owner plus independent critics. This is a direct lesson from Claude of Duty: parallel directory passes increased cross-system defects, while sequential concern ownership, reproducible visual baselines, percentile profiling, and adversarial comparison produced the largest measured improvement. The project still reported 5.05/10 and every blind critic preferred the real reference; the loop requires honest ceilings, not ceremonial claims of perfection. Source: X/@mattshumer_, 2026-07-25

Termination and intervention matrix

Persistent work is bounded by harness state, not by a model promise to stop. Declare the applicable exits before execution and make their counters or event checks observable in the run receipt.

Trigger Harness-owned check Terminal state
Goal met Evaluator passes the frozen acceptance contract complete
Iteration ceiling Completed turns or attempts reach the declared cap exhausted
Spend ceiling Token or currency ledger reaches the declared budget exhausted
Wall-clock deadline Monotonic elapsed time reaches the run deadline expired
No progress The declared state digest or score is unchanged for the configured window stalled
Human interrupt An external kill, pause, or approval signal fires cancelled or pending
Repeated error Consecutive same-class failures reach the retry threshold; any proven success resets the counter failed or pending
External completion A deduped event proves that the PR, ticket, build, payment, or watched condition already resolved superseded or complete
Ambiguity or blocked dependency A consequential fork lacks authority/evidence, or a required parent has not completed pending with an escalation or resume condition

Every non-complete terminal state preserves the last durable checkpoint, reason, budget use, pending authority, retry history, and exact resume condition. Prompts cannot raise their own limits, weaken the evaluator, clear an interrupt, or turn pending into complete. A runtime may expose /goal, /steer, pause, wait, or Kanban controls, but those are interfaces over this state machine rather than substitutes for it. Source: X/@hanakoxbt, 2026-07-20

Execution-event receipt

Long or multi-agent runs should also emit predictable, queryable execution events. Use stable domain.action names, explicit run/goal/task/attempt IDs, primitive privacy-reviewed dimensions, low-cardinality results such as succeeded, failed, retried, cancelled, or exhausted, and one terminal outcome event on every path. Preserve errors and warnings rather than letting a context reducer erase them; keep timing on spans when the observability system supports tracing. A terminal event summarizes the run but never substitutes for its artifacts, state digest, approvals, or deterministic proof. Source: recovered Sentry X Article 2079780939484266496; RTK signal 2059697297626407153, reviewed 2026-08-12

Durable queue and task-brief contract

Minion, fleet, or long-horizon execution is a scheduling topology, not a new workflow truth. Every queued job binds a stable goal and task-brief revision, parent/dependency IDs, input and capability digests, workspace owner, budget, lease/heartbeat, idempotency key, retry ceiling, approval state, output destination, and terminal receipt. Workers may be ephemeral; queue state and proof are durable. A reviewer receives the frozen brief, diff/artifacts, failed checks, warnings, and exact acceptance contract—not a summary produced by the worker being reviewed.

Use a thermonuclear review as a lens only when it is translated into falsifiable checks: unnecessary complexity, avoidable wrappers, leaked domain logic, oversized files, hidden coupling, duplicate abstractions, missing tests, and recovery gaps. Line counts and severity labels are triage signals, not automatic rejection. A task-brief generator may propose milestones and dependencies, but the canonical workflow, authority, evaluator, and completion proof stay in this contract. Source: X 2045427057656729985, 2057521364622553442, 2076902497159987254, reviewed 2026-08-12

Role, worker and capability-scope receipt

A named teammate is a routing and memory identity, not proof of isolation. Every run that crosses agents or persists a worker must record these scopes independently:

  1. authenticated human/service principal and requesting client;
  2. named role and role-instruction revision;
  3. run, attempt, queue item and owning workflow revision;
  4. worker/VM, filesystem/workspace and browser-profile isolation scope;
  5. each connector, login, secret, local-computer permission and network lease, including whether it is role-, worker-, user-, team- or organization-wide;
  6. memory and collaboration scope, including the typed artifact or message used for every handoff;
  7. cancellation, expiry and atomic revocation behavior for the run and every capability; and
  8. deletion/cleanup receipt for schedules, active work, credentials, browser sessions, files, snapshots, transcripts and retained proof.

The UI must disclose shared state before access is granted. A friendly roster, separate chat, or separate screen cannot be labeled isolated when it shares a computer or account login. Keep common workspaces explicit and pass least- capability artifacts between roles by default. Treat model-based Auto Review as contextual steering; deterministic policy and target-bound human approval own the hard boundary. Source: SpaceXAI Grok Bot official overview, computer, skills/routines, approvals/security and team documentation; Cursor Agent and Cloud Agent security documentation, reviewed 2026-08-12

Visual-Proof Receipt

When a visual observation can accept or reject work, the run receipt separates perception from evaluation. Record the baseline and candidate digests, rendered route/state, fixture, viewport and DPR, screenshot hash, selected DOM/scene node or region, and applicable atomic perception axes. Each fact must have a short falsifiable answer and either deterministic corroboration or an explicit reason that only rendered inspection can decide it.

For model-based inspection, also record model and revision, effort, rubric or prompt digest, fallback use, abstentions, repeated-run agreement, and whether an independent critic saw the first answer. A fallback response is attributed to the fallback model, not silently scored as the requested model. Stable aggregate scores do not replace per-case consistency: disagreement across runs or with DOM, data, accessibility, scene, or image-diff evidence leaves the visual fact unresolved.

Failure-derived visual cases should evolve from observed product failures, but the evaluator cannot weaken or relabel its own gate merely because the current candidate fails it. Retain the evaluator digest, case provenance, held-out or public status, license boundary, and full score trajectory. External perception benchmarks may qualify a critic; they do not independently prove taste, interaction, motion, accessibility, or product correctness. Source: PerceptionBench v1, sections 2, 3.5, and limitations; public evaluator at ba032c0

Production-path artifact and readiness proof

When a workflow creates or integrates many assets, components, mappings, or generated states, add these proof types when applicable:

  1. Same-path inspector: a gallery or workbench that imports the production factories, files, materials, animation tables, mappings, and behavior code; record the inspector revision and prove it did not reimplement the target.
  2. Identity matrix: expected object-to-asset, state, action, data, and rights bindings; test present-but-wrong mappings as well as missing or invalid files.
  3. Readiness contract: real milestone events, measurable total or an honest indeterminate state, monotonic state transitions, bounded timeout, error evidence, retry/recovery behavior, and no fabricated percentage.
  4. Specialist-asset manifest: provider/model, input/output hashes, origin, terms/license, attribution, cost, integration path, corrective work, and representative failures for every generated asset class.
  5. Compatibility refusal: the exact supported inputs/platforms and an explicit refusal path when a degraded fallback would misrepresent quality.
  6. Temporal ceiling: capture time and revision for counts, benchmark scores, and quality labels; preserve the strongest known limitation even as the artifact evolves.

These are optional by domain, but once declared they are ordinary frozen proof checks. A prompt, demo, screenshot shelf, loading animation, or successful file parse cannot waive them. Source: X/@0xRishi 2084322235788226653; complete Modern Claudefare live-site, runtime, asset-archive, and media review, 2026-08-12

Persistent Improvement Gate

A workflow that proposes a durable change to instructions, memory, tools, routing, control logic, permissions, skills, workflows, or evaluator rules is a scaffold-update workflow. Its receipt must extend the ordinary completion proof with this state transition:

current digest + learning signal -> candidate patch -> independent verifier
  -> accept next digest | reject and retain current digest

The run must name the updated component, the source or execution signal, the candidate and previous digest, the acceptor, the acceptance suite, and the rollback path. The proposer may run checks, but it cannot be the sole authority that accepts its own structural update. A fresh critic is useful only when its identity, rubric, budget, and evidence are recorded; executable or state-based checks remain preferred when the outcome is mechanically observable.

Evaluator changes are governed more strictly than product changes. Additive or harder cases can enter through the normal review path. Removing, weakening, or rewriting a gate requires separate justification, a visible diff, and approval outside the candidate being judged. Report the complete score/proof trajectory under a declared attempt or resource budget, including regressions and held-out checks; do not publish only the best iteration. Source: Self-Improvements in Modern Agentic Systems: A Survey, sections 6.4, 8, and 9.1; reviewed 2026-08-11

Adaptive Harness Receipt

When a workflow specializes its harness for the current case, the receipt must make adaptation observable rather than hiding it inside a prompt. Record:

  1. the default harness or workflow digest;
  2. the case features visible before execution and the authority envelope;
  3. the retrieved prior runs or global patterns, including both useful failures and successes and their training/evaluation partition;
  4. the exact delta across context, tools, generation, orchestration, memory, and output processing;
  5. a declaration that the current case's hidden label, verifier result, and post-outcome feedback were unavailable during adaptation;
  6. outcome proof plus raw input, cached input, uncached input, output tokens, calls, latency, and price assumptions when cost is compared; and
  7. whether the delta expired, was rejected, or opened a separately gated proposal to change the durable default.

Adaptation is one-shot for the current run. Outcome feedback may be captured afterward for a future training bank, but it cannot revise and rerun the same held-out case while retaining the claim of feedback-free adaptation. Use task correctness as the primary selection criterion; cost is a feasibility gate or tie-breaker. Source: MemoHarness, sections 2.5–3.2, Appendix A–C; reviewed 2026-08-11

Mutation Authority Receipt

Tool execution, durable writeback, and external publication are separate authority surfaces. Approval to run a shell command does not imply permission to commit, push, open a pull request, deploy, send, share, merge, or mutate a third-party service. A saved preference or automation grant may authorize a class of repeated actions, but every run receipt still records:

  1. each enabled mutation capability and the principal or policy that granted it;
  2. target repository, branch, environment, account, audience, and scope;
  3. whether authority was standing, run-specific, or freshly approved;
  4. precondition and expected-head/digest used to prevent stale mutation;
  5. resulting commit, PR, deployment, message, share, or external object ID; and
  6. revocation, rollback, or compensating action.

Defaults fail closed. A workflow may create a feature branch or proposal within an explicit envelope while still requiring separate authority for merge or public release. Open Agents is a useful implementation fixture: dangerous bash commands use one approval policy, while auto commit/push and auto PR are separate saved preferences that default off and use repository-scoped GitHub credentials. Once enabled, however, they can run after a successful turn without a fresh per-turn prompt, so preference gating cannot be mistaken for a universal publish rule. Source: vercel-labs/open-agents at cf865e94, reviewed 2026-08-11

Proactive Trigger Receipt

A prompt, alert, webhook, schedule, anomaly, or saved watch can start work only inside a declared trigger envelope. Record:

  1. trigger kind, source, immutable revision or event ID, observed time, and the correlation or dedupe key;
  2. bound goal/objective, source and object scope, authenticated principal or service identity, and the policy version that admitted it;
  3. allowed capabilities, accounts, repositories, environments, audiences, and the read-only investigation boundary;
  4. time, cost, token, attempt, concurrency, and external-request budgets plus the expiry and kill switch;
  5. idempotency key, dedupe window, expected target digest, and stale-event behavior;
  6. any proposed mutation as a separately hash-bound plan with approver, approval expiry, validation or sandbox requirement, and rollback path;
  7. emitted events, artifacts, decisions, actor/requester/approver attribution, and the final state; and
  8. retry/backoff, escalation owner, dead-letter reason, recovery checkpoint, revocation, and compensating action.

An external alert authorizes observation only unless the standing policy says more. Monitoring an exception, attack, usage spike, or scheduled condition does not silently grant authority to patch code, change configuration, deploy, send, or publish. The built-in product agent and a user's own harness should consume the same underlying tools and skills; only the authenticated identity and authority envelope may differ. Source: X/@rauchg, 2026-07-16; Vercel Agent documentation, reviewed 2026-08-11]

Decision Audit Gate

Tests and diff review do not expose every locally successful but structurally wrong choice. Before merge, the implementing agent must produce a compact decision audit:

  1. list consequential choices, assumptions, shortcuts, and defaults made during the run;
  2. identify choices it is least confident in and the evidence that would falsify each one;
  3. compare those choices with the current owner, acceptance checks, and the general case—not only the fixture that passed;
  4. reopen implementation when a choice patches the example while leaving the underlying failure intact;
  5. preserve the accepted choices and rejected alternatives in the run receipt.

This is choice-surface review, not permission to stop inspecting code. Sample or deep-read code where the decision audit, risk, ownership boundary, or test evidence points. Victor Taelin's MatMul example is the failure fixture: doubling a buffer made the benchmark pass but did not solve the general scheduling problem; asking the agent to enumerate its decisions surfaced the shortcut. Source: X/@VictorTaelin, 2026-07-18

A “pride gate”—asking whether the agent stands behind the branch—is a useful self-critique prompt and can seed this audit, but it is not independent proof. Keep deterministic tests, runtime/visual evidence, and a fresh critic in the gate. Source: X/@vec0zy, 2026-07-18

Generated Closeout

OKF attestation boundary

An OKF v0.2 Attested Computation is a specialized portable computation contract, not a synonym for this run receipt. Use it only when the computation, runtime, declared parameters, executor, and deterministic no-LLM attester are all defined. The agent may fill declared parameter values but must not rewrite the computation. Keep executor receipts outside the knowledge bundle, as the spec requires, and keep concept verified events separate from per-run attestation. Subjective research, design inspection, and ordinary agent runs continue to use this workflow contract without claiming attestation. Source: OKF v0.2 SPEC §6, frozen at knowledge-catalog@374e0bc4

The closeout depends on what changed:

Change Minimum closeout
Handwritten wiki page npx tsx scripts/build-index.ts; qmd update && qmd embed; log when user-visible.
Workflow, resolver, or routing page Page closeout plus npm run routing-doctor; validate affected graph bindings.
Skill source npm run skill-registry; npm run skills:check; routing check; page closeout.
Brain graph or source binding Brain runtime tests; npm run brain; affected workflow checks; page closeout.
Public OKF projection npm run okf:check; verify privacy exclusions and bundle digest; page closeout.
Automation schedule or state model Check freshness and scheduler surface; update state only through the owning run or script.

If a generated surface changes, say which source or generator made it change. Do not make the generated projection the new authority.

Promotion Rule

Workflow output is not automatically wiki truth. Promote when the finding changes durable context:

  • a person, project, tool, decision, workflow, or skill changes
  • repeated friction should become a skill, automation, style rule, or guardrail
  • a source corrects stale compiled truth
  • a run proves an operational state that future agents need

Do not promote routine noise, transient reminders, raw private text, or low-confidence findings. Preserve them in output when useful.

State Discipline

state.json should tell the next agent what is fresh. It must not hide partial failure.

State class Update when
automations.<slug> The scheduled check completed and wrote output/no-op/blocker proof.
syncs.<source> The external source was actually collected or confirmed current.
compile cursors The source was actually reviewed, promoted, or explicitly deferred through that cursor.
generated-surface timestamps The owning generator or index command completed successfully.

For sub-daily jobs, use ISO timestamps. Date-only state hides stale four-hour loops.

Failure Handling

If a workflow cannot finish:

  1. Write the blocker in output or the final answer.
  2. Leave source sync or compile state stale when collection/promotion failed.
  3. Do not claim qmd, generated projections, registries, or redirects are current unless the command ran.
  4. Leave a concrete next action: credential needed, source missing, ambiguity, command failure, or validation gap.

Timeline

  • 2026-08-12 | Added the role/worker/capability-scope receipt after the Grok Bot deep audit: record principal, role, run, VM, workspace/browser, connector/credential/network lease, memory/handoff scope, cancellation and complete revocation separately; do not mistake a named chat or screen for isolation. Source: SpaceXAI Grok Bot official docs; Cursor security docs

  • 2026-08-12 | Added the release/self-update proof profile: source and config identity, prepare/changelog preview, target manifest, artifact hashes, SBOM/provenance/signing, platform smoke tests, separately approved publish digest, remote receipts, authenticated delta chain, bounded resources, full-download fallback, rollback, and primary-source link identity. Source: five-post Sentry release-toolchain replay and five pinned repository receipts

  • 2026-08-12 | Added production-path inspector, semantic identity-matrix, measured readiness, specialist-asset provenance/rights, compatibility refusal, and temporal-ceiling proof profiles from the complete Modern Claudefare follow-on review. Source: X/@0xRishi 2084322235788226653; modernclaudefare.com

  • 2026-08-12 | Added the durable queue/task-brief contract and translated Minions, long-horizon briefs, and thermonuclear review into topology-neutral, machine-checkable job and review requirements. Source: final frontier cohort

  • 2026-08-12 | Added stable execution-event receipts for long and multi-agent runs, including terminal outcomes, trace correlation, privacy-reviewed dimensions, and raw warning/error preservation. Source: Sentry structured-logging article; RTK review

  • 2026-08-11 | Added the proactive-trigger receipt: bind prompt/alert/webhook/ schedule revisions to goal, identity, scope, budgets, idempotency, read-only default, a separately approved mutation digest, audit, retries, dead-letter, revocation, and rollback. Monitoring no longer implies repair authority. Source: X/@rauchg; current Vercel Agent control contract

  • 2026-08-11 | Separated the general workflow receipt from OKF v0.2 Attested Computation: only deterministic executor/attester paths use the portable concept, receipts stay outside the bundle, and verification is not inferred from generation. Source: OKF v0.2 SPEC §6

  • 2026-08-11 | Added the mutation-authority receipt: distinguish tool use, writeback, and publication; record standing versus fresh authority, target, expected digest, external result, and rollback; fail closed by default. Source: X/@rauchg; Open Agents repository authority audit

  • 2026-08-11 | Added the visual-proof receipt: pin rendered state, isolate atomic perception from judgment, attribute fallbacks honestly, require repeated agreement or deterministic corroboration, and preserve evaluator provenance/held-out status. Source: X/@Kimi_Moonshot 2081813202514681878; arXiv:2607.24957v1; public evaluator audit

  • 2026-08-11 | Added the adaptive-harness receipt: base digest, pre-outcome case features, retrieved experience, explicit six-surface delta, leakage declaration, cache-aware cost telemetry, and expire/reject/promote disposition. Source: X/@omarsar0; arXiv:2607.14159v1

  • 2026-08-11 | Added the persistent scaffold-update gate: explicit update target and signal, prior/candidate digests, independent acceptance, rollback, monotone evaluator governance, and full trajectory reporting under a declared budget. Source: X/@omarsar0; arXiv:2607.13104v1

  • 2026-08-10 | Added the proof-driven quality loop: binary checks, baseline, dependency-aware assignment, deterministic and comparative proof, fresh-context criticism, falsification, evaluator hardening, and honest bounded stopping. Source: User request; X/@mattshumer_; mshumer/Claude-of-Duty

  • 2026-08-10 | Added the decision-audit gate: enumerate low-confidence choices, test the general case, retain rejected alternatives, and use a pride prompt only as supplemental self-critique. Source: X/@VictorTaelin; X/@vec0zy

  • 2026-07-01 | Added generated closeout classes and removed the self-link from related pages so this contract cleanly points to proof, state, and generated-surface owners. Source: User request, 2026-07-01

  • 2026-07-01 | Created as the shared proof/state/promotion/validation contract for every workflow page. Source: User request, 2026-07-01