Agentic Quality System
An agentic quality system is the full measurement and verification loop around AI-assisted work: tests, evals, doctors, traces, human review, production monitors, and memory updates.
Why this needs its own page
The wiki has many quality leaves: The Eval Loop (Slop Is an Output Problem), Eval-Driven Development, Agent Trajectory Evaluation, Doctor Pattern, Self-Maintaining Systems, Security and Review Skills, Browser Testing Skills, Frontend and Design Skills, CI checks, browser tests, and wiki doctor. The missing concept was the system boundary: quality is not one check, it is a stack.
Quality stack
| Layer | Catches | Example |
|---|---|---|
| Static checks | syntax, types, lint, formatting, obvious policy violations | TypeScript, Ruff, Ultracite, style doctors |
| Unit/invariant tests | deterministic correctness | small tests around business invariants |
| Browser/runtime tests | user-visible behavior | Playwright, browser-harness, screenshots |
| LLM evals | stochastic output quality | Eval-Driven Development, LLM-as-judge, rubric scores |
| Trajectory evals | path quality and tool behavior | Agent Trajectory Evaluation, traces, tool-call scoring |
| Human review | taste, judgment, risk, final accountability | PR review, design review, approval gates |
| Production monitors | real-world drift, latency, incidents | Self-Maintaining Systems, SLO alerts |
| Memory update | future recurrence | wiki updates, skill changes, new doctor checks |
The loop
The quality system is only complete when failures become future gates. Otherwise the team is just manually inspecting the same class of error forever, which is extremely funny if your goal is suffering.
Agent-specific additions
Agents require extra quality surfaces because they are non-deterministic and tool-using:
- trace the trajectory, not just the final message;
- score tool choice and tool arguments;
- preserve artifacts and state writes;
- separate evals for output quality, safety, and cost;
- make approvals part of the run record;
- update memory only after evidence is reviewed;
- encode repeated failures into skills/doctors.
Compose gates without hiding evidence
Fast static tooling is valuable when it shortens the feedback cycle without changing what can be seen. TypeScript 7/tsgo, Oxlint/Oxfmt, type-aware linting, Impeccable's deterministic checks, and similar combined passes should be tested on representative repositories against the incumbent compiler, formatter, linter, and tests. Record compatibility, diagnostics gained and lost, false positives, migration work, warm/cold latency, changed-line latency, and the exact command/version. A faster aggregate command is a front door, not a reason to delete distinct semantic gates.
The same rule applies to context reducers such as RTK and documentation shells
such as agentsearch.sh: preserve command, exit status, warnings, errors,
provenance, freshness, truncation, authentication, and a raw-output escape
hatch. Measure task success and debugging recovery alongside token savings.
For long runs, emit stable domain.action events with low-cardinality terminal
results and a final success/failure/cancelled event; keep privacy-approved
dimensions explicit rather than logging arbitrary objects. Source: X Articles
2039817244696567808 and 2079780939484266496; X 2031194113413148747,
2074899956511760806, 2079774630391226859, 2059697297626407153;
current TypeScript and RTK repository evidence, reviewed 2026-08-12
Build the verifier surface before scaling the loop. Each worktree or sandbox
should expose a reproducible app instance, DOM/screenshots/video, logs, metrics,
traces, deterministic commands, and repository-local acceptance criteria to the
agent. Superpowers' design gate and staged spec/code review, Cursor's strict
maintainability skill, and cross-harness reviewers are useful procedures only
when their findings are checked against the target product and deterministic
gates; harshness, a different model family, or more reviewer personas does not
by itself make a critic independent or correct. Source: OpenAI harness
engineering report; obra/superpowers@44c9b2d; cursor/plugins@6dbbdd5,
reviewed 2026-08-12
Coupling-Aware Improvement
An agentic quality system must model dependencies before it fans work out. Independent files or fixtures can progress in parallel. Coupled concerns—such as lighting, tonemapping, materials, and sky—need one sequential owner because a local improvement can invalidate the assumptions of every neighboring pass. Independent critics should remain separate from implementation ownership.
Claude of Duty supplies measured evidence: three parallel six-agent rounds left
cross-system visual defects higher, while a sequential concern-owned pass moved
the critic score by a full point and cut defects. Its reproducible baseline tool
also proved that reused-page screenshot sets were nondeterministic, and its
gameplay profiler showed a static median hiding 728–1236 ms stalls. Quality
therefore needs isolated baselines, percentile/tail metrics, and falsification of
the brief itself—not merely more workers or more iterations. Source: mshumer/Claude-of-Duty README and ARCHITECTURE.md, commit d9b237b, checked
2026-08-10
Cursor's SQLite swarm reinforces the boundary from the other direction: parallelism can work when a planner owns an explicit task tree, workers receive bounded context, a neutral reconciler owns merges, and independent critics use different review lenses. Measure context efficiency, duplicate work, merge conflicts, worker-token share, time, cost, and held-out quality; agent count is not itself a quality layer. Source: Cursor, “Agent swarm model economics,” 2026-08-07
Production-Path Inspection And Domain Correctness
For products with many generated assets, behaviors, or states, prefer an inspection surface that imports the same factories, files, materials, animation tables, mappings, and behavior code as production. A parallel demo gallery can look healthy while the shipped path drifts. The inspector should let an agent or human isolate one item, compare variants, inspect provenance, and exercise its real integration without cloning the implementation.
Presence and runtime health are only the first gates. Add a semantic identity matrix for valid-but-wrong outcomes: the expected weapon-to-audio mapping, model-to-rig/animation binding, icon-to-action meaning, data-to-chart series, component-to-state behavior, or asset-to-license record. A file that loads can still belong to the wrong object; a procedural mesh that renders can still be disconnected or mechanically impossible.
Readiness UI is proof-bearing too. Bind visible milestones to observable runtime events, use an indeterminate state when the total is not measurable, and enforce a bounded timeout with a diagnosable failure state. Never animate a percentage merely because work is occurring. For specialist-generated assets, record provider/model, exact input and output hashes, cost, origin, license or terms, attribution, integration path, corrective work, and representative failure. If the target platform or input is unsupported, refuse explicitly rather than presenting a degraded fallback as equivalent.
Keep the ceiling honest and quantitative claims time-stamped. Modern
Claudefare's launch author explicitly called the impressive browser prototype
non-AAA, while the live build later moved from 20 guns and 430-plus audio files
to 23 weapons and 965 recordings. Iteration changed the artifact; it did not
retroactively prove the stronger label. Source: X/@0xRishi
2084322235788226653; captured Modern Claudefare About, Asset Archive,
runtime, and 137.856-second launch artifact, 2026-08-12
Repeated project evidence is a reusable quality contract
The agent-session corpus confirms the same contract across unrelated products. Agent-Pets evaluates every direction and animation, preserves blind-review disagreement, and distinguishes metric warnings from visible failures. Claude of Tanks measures each asset across fixed views, zooms, masks, sub-scores, and two identical floor runs while checking sibling assets for regression. Princeton Tower Defense couples a stress scene and p95 budget to a build-blocking balance audit. Glyphfield keeps one production renderer, invalidates only the affected pass, and tests rapid switching and export rather than trusting a demo.
General rule: freeze the identity/state/view matrix, run the production path, record a deterministic baseline, score both the minimum floor and tail behavior, separate the builder from the critic, and preserve disagreement or no-output as evidence. A high average may not hide one broken direction, viewport, asset, animation, or long-frame stall. A failed generation, missing artifact, or indeterminate metric is a first-class result—not a sample to drop until the chart looks complete. Source: reviewed Agent-Pets, Claude of Tanks, Princeton Tower Defense, Glyphfield, Sandbox Arena, and Sigil UI agent-session lineages, 2026-08-12
Codex as feature tester
@gdb's Codex testing bookmark is a useful product signal: coding agents are moving from "write this patch" toward "exercise every feature path in the app and produce evidence." Route that work through the quality stack, not through ad-hoc clicking: the agent can explore, use browser/runtime checks, capture screenshots and traces, then promote important paths into Playwright, invariants, or evals. Source: X/@gdb, 2026-06-21
Production-scale quality is trust per decision
DoorDash's Flux announcement reports more than 25,000 code reviews each week,
while its earlier code-review writeup makes the important acceptance metric
behavioral: whether engineers act on comments and keep the reviewer enabled.
At that volume, raw finding count is actively dangerous. Track accepted changes,
false-positive and duplicate-comment rates, coverage, missed-risk audits,
latency/cost, repository and rule drift, escalation, disablement, and rollback.
The local code.pr-review manifest already encodes coverage, evidence,
falsification, and gated learning; future scale work should preserve those proof
objects rather than optimizing for reviews-per-week alone. Source: X/@AIatDoorDash; DoorDash code-review engineering writeup
Kevin rule
When a page says "verify," it should name the quality layer:
- deterministic code: tests/types/lint,
- UI: browser screenshot and interaction checks,
- agent behavior: trajectory/eval suite,
- content/taste: rubric + humanizer/voice skill,
- wiki: build-index, doctor, qmd, log,
- production: monitors and rollback path.
Vague verification is just vibes with a clipboard.
Timeline
-
2026-08-12 | Folded cross-project agent-session evidence into one reusable quality contract: identity/state/view matrices, same-production-path inspection, deterministic baselines, minimum floors plus tail metrics, sibling-regression checks, critic separation, and explicit no-output states. Source: contract-v4 agent-session lineage review
-
2026-08-12 | Added DoorDash Flux and its adjacent production reviewer as scale evidence: optimize trust per decision, not finding volume, and preserve coverage, falsification, accepted-change, false-positive, duplicate, cost, and rollback proof when
code.pr-reviewmoves to remote worker fleets. Source: X/@AIatDoorDash2087285008906240193; DoorDash engineering -
2026-08-12 | Added same-production-path inspectors, semantic identity matrices, measured-versus-indeterminate readiness, specialist-asset provenance/rights receipts, unsupported-input refusal, and time-stamped ceiling claims after a complete Modern Claudefare site/runtime/media review. Source: X/@0xRishi
2084322235788226653; modernclaudefare.com -
2026-08-12 | Added gate-composition and evidence-preservation rules for fast TypeScript/lint stacks, output reducers, documentation shells, and structured terminal events. Speed is accepted only with diagnostic and task parity plus a raw-evidence escape hatch. Source: quality/skills cohort
-
2026-08-12 | Added verifier-surface-first routing from OpenAI's primary harness report and current Superpowers/Cursor skill sources: agents need worktree-local app, browser, observability, commands, and acceptance evidence before loops scale; reviewer count or harsh phrasing is not independence. Source: OpenAI harness engineering; obra/superpowers; cursor/plugins
-
2026-08-10 | Added coupling-aware assignment, isolated reproducible baselines, tail-performance checks, independent critics, and brief falsification from the Claude of Duty postmortem. Source: User request; X/@mattshumer_; mshumer/Claude-of-Duty
-
2026-07-04 | Added Codex-as-feature-tester as a quality-system signal: agents should explore app behavior and produce evidence, but recurring paths must be promoted into browser tests, invariants, or evals. Source: X/@gdb, 2026-06-21
-
2026-06-19 | Created to unify evals, doctors, traces, CI, reviews, production monitors, and memory updates into one cross-section quality concept. Source: whole-wiki concept review, 2026-06-19