The Eval Loop (Slop Is an Output Problem)

Slop is an output problem, not a prompt problem. Better prompts, bigger models, memory, and context files are all input-side fixes for a thing that was never broken. The missing layer is measurement: an eval loop that scores every output against a standard before (and after) it ships. Almost nobody runs this loop — that's the slop.

The eval loop: slop is an output problem, not a prompt problem

The reviewed Hermes-operator artifact independently compresses this page's thesis: "slop is an output problem, not a prompt problem." Its diagram contrasts input-side tweaks (better prompts, bigger model, memory, context files) with the missing output gate: score every run 0-1 against a benchmark, ship only above threshold, and feed failures/thumbs-down back into the case set. Treat it as corroborating operator evidence for the existing Hermes Agent (Nous Research) eval-loop route, not a separate new framework. Source: X/@EXM7777 and local artifact review, 2026-07-03

The core argument

A prompt is a hypothesis, the output is the result, and an eval is the only thing that closes the loop between them. Because the model is non-deterministic, the same prompt run twice gives two answers — so even a perfect prompt produces slop on some fraction of runs, and you have no idea which runs until a user is staring at one. A perfect prompt isn't a quality guarantee; it's "a slightly better coin flip, and you're shipping every flip." Source: User essay, 2026-06-01

People keep reaching for prompts because the prompt is the only lever they can see and editing it feels like control. Measurement is invisible and unsexy, so the whole conversation stays stuck on the one lever that can't solve the problem alone. The people whose AI output is consistently clean aren't better at prompting — they have a second lever: they measure every output against a standard before it ships. Source: User essay, 2026-06-01

The analogy: a factory shipping defects doesn't have a worker problem, it has a quality-control problem — nobody checks the output before it leaves the building. You built half a system (the generate half) and blamed the half that works.

Two places slop lives — same disease, same cure

  • Content output — tweets, articles, emails, posts. Slop looks technically fine and completely hollow; it dies loudly in public.
  • Product output — the AI feature, agent, chatbot, extraction pipeline. Slop is a confident wrong answer, a hallucinated number, a broken JSON payload, a tone drift three deploys later; it scales in silence as users quietly leave.

Both are un-measured AI output going straight to an audience with no gate in between. The only difference is stakes and visibility. One quality system grades both.

The eval loop = unit tests for the non-deterministic

You are not testing whether the code runs; you are testing whether the output is good, across enough cases that one bad run can't hide. It runs in three places:

  1. Before you ship — run the new prompt/model against a saved case set and confirm it didn't get worse (regression testing; stops a change that fixes one thing and silently breaks three).
  2. At runtime — score output as it's generated and let conditional logic catch failures before the user sees them (guardrail).
  3. In production — score a sample of real executions continuously, so you see quality degrading the day it starts, not the week a client complains.

The moment quality becomes a number, slop stops being a feeling and becomes a bug. "You can't debug a vibe; you can debug a score that dropped from 0.82 to 0.61."

Scientific experiment branch

Model, harness, retrieval, and other empirical optimization needs a stricter version of the same loop. Freeze the baseline, data/evaluator revisions, primary and guardrail metrics, budget, seeds, and held-out set before changing the candidate. Research prior recipes and bind their claims to exact methods; audit the data; preflight the exact source in a representative sandbox; run one pilot before a batch; then diagnose and ablate one causal variable where possible. Keep a candidate only when it beats the frozen baseline without violating guardrails and survives untouched held-out confirmation. Source: huggingface/ml-intern system and research prompts, reviewed 2026-08-10

Use Research Experiment Workflow for the complete, resumable process. It adds authority, cost, data-rights, trace-egress, provenance, and exact-artifact proof that a generic output eval does not need. ML Intern - Hugging Face Research Agent is one possible executor for HF-native ML work; AutoResearch (Karpathy) - Applied to Marketing is the smaller fixed-scorer code ratchet. Neither tool owns the scientific claim. A launch chart or best-of-many run remains a hypothesis until the receipt exposes data, code, evaluator, selection policy, costs, and replay evidence.

Use an evidence ladder so high-signal sources can be retained without promoting their strongest framing. A social demo or one-shot supplies a candidate and failure hypotheses; a primary article or repository establishes what was actually built; a public leaderboard describes only its submitted condition; an owned representative run establishes local behavior; and untouched held-out cases plus independent acceptance gates justify a default change. Missing a lower rung does not make the source useless—it caps the claim. This applies to model arenas, planner/worker/judge recipes, strict review skills, visual clones, generative-media demos, and large agent fleets alike.

A private production benchmark can be highly actionable without being externally reproducible. Ramp SWE-Bench uses 80 tasks reconstructed from merged production PRs, a fixed base commit, held-out solutions and tests, single pass@1, and a 45-minute timeout. That is strong evidence for the method and for Ramp's internal routing decisions. Because prompts, patches, tests, and repository states remain private, it is not independent proof that another harness will reproduce the aggregate score. Preserve both truths: learn from the benchmark construction, but require an owned representative suite before changing Kevin's default model or harness. Source: X/@RampLabs 2065485806605619304; Ramp SWE-Bench engineering write-up, reviewed 2026-08-12

Skill and comprehension experiments

Treat a skill change as an experiment over both routing and behavior. Freeze the current skill digest, trigger set, representative tasks, budgets, evaluator, and holdout before proposing an edit. Separate the inner task loop from the outer learning loop: completed runs and human corrections may propose a small skill/helper/rubric change, but a distinct acceptor compares old and new on the same cases, per-skill regressions, false triggers, cost, stagnation, and rollback. Acceptance, praise, or comment volume is feedback—not ground truth. Role-diverse reviewers count as independent only to the extent their evidence, prompts, models, tools, and failure modes actually differ. Source: recovered X Articles 2066906439407276032, 2071663234772574209, 2077369329025196032, and 2077425822084759552; warpdotdev-demos/cloud-factory-demo@bd3a79ae, reviewed 2026-08-12

For teaching and explanation systems, evaluate more than whether the learner liked the artifact. Compare a written explanation, targeted quiz, interactive microworld, oral defense, and novel problem-solving transfer on the same concept. Record source accuracy, correction quality, time, accessibility, anxiety, retention, false confidence, and cost. Use the least burdensome mode that diagnoses the failure; no single modality is a universal comprehension oracle. Source: Hacker News 49224294, 49224497, 49224973, 49234675, 49239268; reviewed 2026-08-12

The same evidence ladder applies to efficiency and hardware claims. Report model/checkpoint and quantization, weights and KV cache, runtime and kernels, prompt/context/output lengths, batch/concurrency, warmup, prefill and decode rates, peak and steady-state RAM/VRAM, shared-process accounting, hardware, power, price date, quality delta, failures, and rebuild/update behavior. A process screenshot, headline parameter count, compressed index size, or best-case tokens/second cannot establish local feasibility or total cost.

Repository structure is evidence, not performance by itself. Current retained examples illustrate distinct hypotheses: ramp-public/portallib@1a7b8c5b (Apache-2.0) for portable task adapters; MoonshotAI/MoonEP@7745ffa0 (MIT) for expert-parallel communication; kvcache-ai/ktransformers@eb9b70c4 (Apache-2.0, 146 test files) for heterogeneous local inference; and google-research/google-research@01553912 (Apache-2.0) as the upstream research tree containing vector-search work. Each still needs an exact method/model/data slice and representative owned run. Source: commit-pinned repository evidence, captured 2026-08-12

Benchmark breadth versus regression depth

Do not make one suite serve two incompatible jobs. A benchmark suite spans enough representative tasks, models, harnesses, and environments to compare candidate conditions. A regression suite is smaller, faster, and deeper: it preserves known failures, edge cases, and release blockers for the selected condition. Both need explicit scenario revisions, attempts, timeouts, terminal checks, and cost, but their sampling and cadence differ.

Supabase Evals demonstrates the useful shape: scenarios pair a PROMPT.md with an executable EVAL.ts; experiments bind an agent/runtime/model condition; tool-mode scenarios call a Management-API-compatible platform-lite backed by PGlite; local-stack scenarios exercise the Supabase CLI and Docker stack; and scorers inspect copied workspaces and live state rather than trusting process exit zero or final prose. Its parser tests normalize Claude, Codex, and OpenCode transcripts and preserve failed turns. One retry may measure recovery, but the first failure and retry must be reported separately so retry does not erase reliability. Source: supabase/evals@5f8cfeb9, README, workflows, scenario evaluators, transcript parsers, and runner tests; Supabase Evals launch article, reviewed 2026-08-12

"Real tools" is not one environment class. Every result declares whether the target is an emulated API, a lightweight local implementation, a real local stack, or a hosted service; declares network, credentials, Docker/socket, ports, seeded state, and cleanup; and states which integration checks were actually enabled. platform-lite is valuable because the agent uses real MCP calls against an API-compatible surface, but it is not hosted Supabase. Docker integration tests gated behind SANDBOX_DOCKER_TESTS=1 are not proven by the default check. Track skill activation and documentation use as mechanisms; promote only on terminal task/state outcomes, adversarial cases, and held-out regressions.

The benchmark: three parts (content and product)

Part Content Product
Test cases (ground truth) 20–50 of your best pieces (the gold standard you already hit on your best days) Real inputs pulled from logs / user sessions — the weird ones, not the happy path
Metric (output → 0–1) A specific rubric (e.g. "actionable / followable / replicable / novel"; meta-criterion: would someone bookmark and implement this?) Match the task: exact match (one right label), validator (structure must hold), semantic similarity + judge (open-ended)
Threshold (the line) 0.7 to start; nothing below ships 0.7 to start; below baseline blocks the deploy

Skip any one part and you have a wish, not a gate. A vague rubric ("is this good and engaging") produces a vague score; a specific rubric produces a score you can trust — "the judge inherits your taste only if you actually write your taste down." The threshold only works if you never let a 0.6 through because you liked it (takes the late-night ego out of the decision). Source: User essay, 2026-06-01

BINEVAL is the current concrete method for making LLM-as-judge rubrics inspectable: break a criterion into atomic yes/no questions, answer them independently, and aggregate the verdicts. Use it when a holistic judge score hides why an output failed; keep deterministic validators first where the task is symbolic. Source: arXiv:2606.27226, 2026-06-25

Diana Hu (YC) compresses the same point to a slogan: "taste is just an eval you haven't written down yet." You make taste legible by reading 1,000+ real failure traces, grouping them into a failure taxonomy (what researchers call open / axial coding), and treating that taxonomy as the eval suite. Generic benchmarks (MMLU, HLE) measure "is it smart"; you need to measure "did it do the job" — different questions. Because failure modes are domain-specific, "evals are the moat now, not code" (same models, leaked prompts, scrapeable data), and that is precisely why vertical AI companies can beat general models. A benchmark you can saturate but can't tie to customer value is a vanity metric. Source: X/@sdianahu, 2026-06-04

The compounding arrow (the whole game)

Failures + thumbs-down on production runs become new test cases written back into the suite. That failed run becomes a permanent check, so the quality floor rises on its own every week. In the diagram this is the dashed arrow from "failures + thumbs-down" back to "benchmark" — "this arrow is the whole game," and almost nobody runs it.

Red Queen Gödel Machine adds the next constraint for agent self-improvement: the evaluator cannot stay frozen forever, but it must stay frozen inside each epoch. Failures and saturation produce evaluator challengers—more precise deterministic checks, harder held-out cases, or BINEVAL-style atomic questions—without moving the active target. At a declared checkpoint, a separate acceptor compares challenger and incumbent on the same evaluator-independent anchor. A promoted evaluator invalidates the old evaluator's dependent scores before re-ranking; unanchored winners remain epoch-local. The implementation route is Agent Self-Improvement Eval Library, and Loopy owns the bounded loop around the evaluator. Source: arXiv:2606.26294v2, replayed 2026-08-11

Why it matters here

This is the same gate-shape as teacher-mode (verify against a standard, fail → re-teach, completion gate) and the verification-before-completion ethos — evidence before assertions, applied to output quality. It is the cure for the vibe-driven failure mode, and it is what Hermes is built to run as a standing loop (skills as the judge, memory as ground truth, cron as the monitor, approval buttons as the gate). For Kevin's own content, the Writing and Content Skills + kevin-voice skills are the content rubric/judge — the eval loop applied to his writing. For user-facing apps that must block harmful text/images, OpenAI Moderation API is a complementary runtime harm classifier (policy signals); it does not replace custom rubrics for product quality. Source: compiled, 2026-06-01; updated 2026-06-12


Timeline

  • 2026-08-12 | Added controlled skill-evolution and comprehension-evaluation branches: frozen digests and holdouts, separate proposer/acceptor, real independence checks, and multi-modal transfer rather than satisfaction-only scoring. Source: quality/skills/HN cohort

  • 2026-08-12 | Added a component-aware efficiency/hardware receipt and retained commit-pinned PorTAL, MoonEP, KTransformers, and Google Research sources as distinct experiment hypotheses. Headline compression, parameter, memory, latency, and price claims no longer promote without exact model/data/runtime/hardware/quality/failure proof. Source: model, memory, hardware, and HN source cohort; repository evidence

  • 2026-08-12 | Added the private-benchmark boundary using Ramp SWE-Bench: production-grounded, held-out internal tasks can be strong organizational evidence while private prompts/tests/repository states still cap external reproducibility and default-routing authority. Source: X/@RampLabs 2065485806605619304; Ramp engineering write-up

  • 2026-08-12 | Added the evidence ladder for high-signal launch posts: retain the candidate, but cap claims at social artifact, primary source, public benchmark, owned representative run, or held-out proof. Reconciled current agent-loop, visual-model, review-skill, retrieval, generative-media, and security-fleet signals without turning one-shots or leaderboards into defaults. Source: 42 saved X records; OpenAI harness engineering; Cursor MiniSQLite; obra/superpowers; Ramp security engineering; Elastic docs

  • 2026-08-11 | Corrected the Red Queen rule from continuous hardening to controlled utility evolution: freeze evaluator and protocol inside an epoch, propose challengers from failures, promote only at anchored independent checkpoints, invalidate displaced-evaluator scores, and scope unanchored comparisons to one epoch. Source: arXiv:2606.26294v2; CaMLSys explainer, 2026-07-19

  • 2026-08-10 | Added the scientific experiment branch and routed it through the canonical research-experiment workflow: frozen baseline and evaluator, attributed recipes, data audit, exact-source pilot, causal ablation, held-out confirmation, and provenance-bearing writeback. Source: X/@akseljoonas; huggingface/ml-intern, reviewed 2026-08-10

  • 2026-07-03 | Deep-reviewed the Hermes anti-slop operator diagram and attached it as artifact evidence for the same eval-loop thesis: input-side prompt/model/context changes are half a system without an output score gate and failure-feedback loop. Source: raw/x-bookmarks/enriched/2061086049326256356.json

  • 2026-06-29 | Added the Red Queen evaluator-evolution rule and linked the local self-improvement eval runner as the implementation route. Source: arXiv:2606.26294; User request, 2026-06-29

  • 2026-06-26 | Linked Awesome Evals as the source map for agent-eval papers, talks, tools, benchmarks, LLM-as-judge patterns, trajectory grading, and RL-environment thinking. Source: GitHub benchflow-ai/awesome-evals, 2026-06-26

  • 2026-06-28 | Added BINEVAL as the binary-question method for turning vague LLM-as-judge rubrics into inspectable criterion-level verdicts. Source: arXiv:2606.27226

  • 2026-06-18 | Added Agent Trajectory Evaluation as the agent-run extension of this output-quality loop. Source: whole-wiki graph audit, 2026-06-18

  • 2026-06-15 | Phase 7 (Review) mapped to the eval loop: Matt Pocock's review skill checks Standards (rubric) + Spec (matches the PRD) in parallel sub-agents, and tdd's red-green-refactor is the same gate at the unit level. Source: github.com/mattpocock/skills, 2026-06-15

  • 2026-06-15 | Satya Nadella's essay "A frontier without an ecosystem is not stable" corroborates the thesis from firm economics: "private evals should capture whether a model is actually improving against outcomes that matter to the business (not just external benchmarks)," plus private RL on real internal traces — i.e., the eval loop + compounding arrow as the firm's IP. See Own the Learning Loop. Source: X/@satyanadella, https://x.com/satyanadella/status/2066182223213293753; surfaced by User, 2026-06-15

  • 2026-06-12 | Cross-linked OpenAI Moderation API as complementary harm-classifier for user-facing runtime guardrails (distinct from custom eval rubrics). Source: User capture, 2026-06-12

  • 2026-06-04 | Diana Hu (YC) thread corroborates the thesis: "taste is just an eval you haven't written down yet"; evals come from error analysis (1,000+ traces → failure taxonomy via open/axial coding), are domain-specific, and are "the moat now, not code." 624 likes, 566 bookmarks. Source: X/@sdianahu, 2026-06-04

  • 2026-06-01 | Page created from Kevin's essay "How To Fix AI Slop (Using Hermes)." Compiled the output-vs-input thesis, the non-determinism argument ("a perfect prompt is a slightly better coin flip"), content vs product slop as one disease, the three run-sites (pre-ship regression / runtime guardrail / production sampling), the three-part benchmark, and the failures→new-test-cases compounding arrow. Diagram saved to wiki/assets/eval-loop-diagram.png. Source: User, 2026-06-01

61 pages link here

Active Stack (What to Actually Use)MetaAgent Guardrails ArchitectureArchitectureAgent Harness CapsulesConceptsAgent Harness, Runtime, Memory, and EvalsConceptsAgent LoopingConceptsAgent MachinesProjectsAgent Operations HubMetaAgent Self-Improvement Eval LibraryConceptsAgent Trajectory EvaluationConceptsAgentic Quality SystemConceptsAI BerkshireToolsAutoResearch (Karpathy) - Applied to MarketingToolsAwesome EvalsToolsBeta DX Walk (Friction Is a Bug)ConceptsBINEVALConceptsBorrowed Intelligence (Drain It into Plans)ConceptsBrain CapsulesMetaCapability Routing MapMetaClaude Fable 5ToolsClaude Managed AgentsToolsConcept System MapConceptsContent Pipeline WorkflowWorkflowsContentbitToolsExa AgentToolsFreeLLMAPIToolsHarness EngineeringConceptsHermes Agent (Nous Research)ToolsHermes Harness (Nous Research)ArchitectureHyperagentToolsKimi K3ToolsLanguage OrchestrationResearchLLM CouncilConceptsLLM Synthetic Panels (Purchase Intent)ConceptsLoopyToolsMatt Pocock's Skills (Skills for Real Engineers)ToolsMiniMax M3ToolsML Intern - Hugging Face Research AgentToolsMulti-Agent TopologiesConceptsOpenAI Moderation APIToolsOpenRouterToolsOpenRouter FusionToolsOwn the Learning LoopConceptsRed Queen Gödel MachineConceptsRed-Team Your Business (Adversarial AI Audit)ConceptsReflection and Self-Critique Agent ArchitectureArchitectureResearch Experiment WorkflowWorkflowsSatya NadellaPeopleSkill ResolverSkill ResolverSocial Content OSToolsStaying in the Loop with AgentsConceptsSuperpowersToolsTaste as a MoatPhilosophiesThe 7 Phases of AI-Powered DevelopmentWorkflowsThe Brain-Agent LoopConceptsTheta SoftwareProjectsVibecoding CritiqueConceptsWiki Concept Coverage MapMetaX Bookmark Ingest 2026-06-28ResearchX Bookmark Ingest 2026-07-06ResearchX Bookmarks: AI Agents & Tools (Jan 2025 – Jun 2026)ToolsX Bookmarks: Design & UI (Feb–Jun 2026)Design