Controlled Skill Evolution - GEPA and SkillOpt

Skills may improve from execution evidence, but only through a controlled candidate-selection-promotion loop. GEPA explores textual variants through Genetic-Pareto search; SkillOpt trains one compact skill through bounded patches and strict held-out selection. Neither licenses unconditional self-rewrite. [Sources: GEPA; SkillOpt]

The Problem It Solves

In-agent skill learning has a known weakness: agents tend toward self-congratulation. They almost always rate their own performance highly even when they failed. The same system that auto-generates skills can overwrite manual customizations with worse versions. Community feedback confirms this is a real and persistent failure mode. Source: compiled from X/@akshay_pachaar article, 2026-05-13

How It Works

  1. Read current skill from the agent's skill directory
  2. Generate evaluation dataset (synthetic via Claude Opus, real session history from SQLite, or hand-curated golden sets)
  3. Run GEPA optimizer: read execution traces -> understand failure points -> generate candidate variants through evolutionary search
  4. Evaluate candidates using LLM-as-judge scoring with rubrics (not binary pass/fail)
  5. Apply constraint gates: full test suite passes 100%, skills stay under 15KB, caching compatibility preserved, semantic purpose doesn't drift
  6. Best variant goes out as a PR - never a direct commit

Cost: $2-10 per optimization run. No GPU required - everything through API calls.

Relevance to Kevin's System

Kevin's skill-maintenance automation already runs periodic skill updates from session analysis. GEPA offers a more rigorous alternative: instead of the agent deciding "I should improve this skill," execution traces objectively show where skills failed, and evolutionary search finds better variants with constraint preservation.

The PR-based output (never direct commit) aligns with Kevin's code review conventions. The self-congratulation problem is exactly why Kevin uses Security and Review Skills - a second model catches what the first model thinks is fine.

GEPA is positioned as a middle ground: more rigorous than in-agent self-improvement, cheaper than full RL fine-tuning (GRPO). The paper's current arXiv v2 says GEPA reflects over trajectories, proposes prompt updates, combines complementary lessons from the Pareto frontier, outperforms GRPO by 6% on average and up to 20%, and uses up to 35x fewer rollouts. Source: arXiv:2507.19457v2, 2026-02-14

SkillOpt sharpens the same skill-evolution problem from the optimizer side. Its diagram treats a skill document as the trainable state of a frozen agent: rollouts produce trajectories, a separate optimizer proposes bounded add/delete/replace edits, a held-out gate accepts only strict validation improvements, and rejected side updates become negative feedback. This is closer to text-space gradient descent than GEPA's broader Genetic-Pareto prompt evolution, but both reinforce the same operational rule for Kevin's system: optimize skills only through trace evidence and validation gates, never through unconditional self-rewrite. Source: X/@muratcan and local image artifact, replayed 2026-08-11; Source: SkillOpt project page

Controlled Skill Optimization Contract

Kevin-Wiki implements the shared lesson as a stricter local contract:

The machine-readable form is skillOptimization in evals/agent-self-improvement/suite.json; eleven contract tests now cover it and the controlled-evaluator contract together. The target model, execution harness, evaluator, and split revisions stay frozen while one skill candidate is compared. Training creates candidates; validation selects; final test only reports. Ties reject, holdout leakage abstains, and the canonical skill cannot change before the gate. Fast patches may not overwrite protected slow policy. Open-ended writing, design, and strategy updates remain held without a reliable verifier or reviewed evidence. Source: local contract and tests, implemented 2026-08-11

Two independently failing surfaces

A skill has at least two measured surfaces:

Surface What the agent sees Required evidence
Router description Frontmatter names/descriptions only top-1/top-k routing, false triggers, near-misses, and per-skill confusion
Activated body Full selected SKILL.md and resources task/state outcome, regressions, cost, and negative transfer
End-to-end Normal runtime discovery and activation the intended description selects the intended body and improves the task

The v2.3.0 Context Engineering Agent Skills benchmark proves why those lanes cannot be merged. Its router sweep deliberately used settingSources: [], so the model saw descriptions but no bodies. Targeted description changes moved context-fundamentals by +23.4 percentage points, project-development by +25, and tool-design by +7.8, while whole-model top-1 deltas were only +2.5, +2.7, +3.9, and -2.0 points across the four reported models. The repository explicitly labels body-loaded effectiveness as a later stage that had not yet run. Aggregate accuracy would have hidden the important skill-level effects; router accuracy would have overstated what had actually been proven. Source: v2.3.0 release, README, changelog, router runner, and published 600-run report; captured 2026-08-11

What SkillOpt Actually Shows

The arXiv v2 protocol uses four epochs, a default rollout batch of 40, reflection minibatches of 8, and a textual edit budget of 4 with cosine decay to 2. Those are experiment defaults, not universal constants. The learning-rate sweep has mixed winners across SearchQA, SpreadsheetBench, and LiveMath; it does not establish a general “4–8 edits is the sweet spot.” Removing the budget made the tested scores worse, not universally collapse. The protected slow/meta ablation's largest drop was SpreadsheetBench 77.5 to 55.0 (-22.5 points); that is a benchmark-specific result, not a universal 22-point law. Source: SkillOpt arXiv v2, Ablations

The paper's learned GPT-5.5 artifacts were 379–1,995 tokens with a roughly 920-token median and only 1–4 accepted updates. That is evidence for compact, selective artifacts, not a target length or edit count that every skill must copy. The clearest portability result is SpreadsheetBench: Codex-trained to Claude Code improved 22.1 to 81.8 (+59.7), while Claude-Code-trained to Codex improved 27.5 to 71.1 (+43.6). LiveMath cross-harness gains were smaller (+1.6 and +12.8), so every destination still needs its own transfer receipt. The authors report best or tied-best results on 52/52 evaluated cells; that claim is bounded to their models, benchmarks, scorers, splits, and harness adapters.

The method depends on scored trajectories and a disjoint selection split. Open-ended subjective work needs stronger human or model evaluation; training adds rollout and optimizer cost; one compact skill can be insufficient for a heterogeneous domain; and transfer still requires careful held-out testing. The deployed skill has no optimizer call, but “zero inference-time cost” means no extra model call—not zero prompt tokens. Source: SkillOpt arXiv v2, Experiments, Limitations, and Conclusion

Adoption Boundary

Borrow Do not infer or adopt
bounded, provenance-bearing patches autonomous canonical rewrites after a single success
strict disjoint held-out improvement; ties reject aggregate-only promotion or leaked holdouts with a warning
rejected-edit ledger and protected slow state treating all rejected edits as bad when the gate itself leaked
separate router, body, and end-to-end lanes calling description accuracy skill effectiveness
per-skill effects, regressions, and transfer checks universal edit-count, token-length, or cross-harness guarantees
exact pinned source as research/reference input installing SkillOpt-Sleep globally or authorizing provider spend

The current SkillOpt main branch adds useful holdout-leak abstention tests and append-only evidence machinery beyond the stable v0.2.0 tag. It also contains an optional gate bonus based on “semantic density” of words such as MUST, ALWAYS, and CRITICAL. Kevin-Wiki does not adopt that bonus: instruction-word density is a gameable textual proxy and cannot substitute for routing or task outcomes. This is a local inference from the inspected current source, not a claim by the paper. The repository remains a pinned research reference rather than a globally installed self-modifier.

Session Capture Is Not Skill Promotion

SkillClaw adds a useful upstream stage: it can intercept model/tool trajectories, collect repeated strategies across sessions, summarize candidate skills, and optionally distribute them through a shared registry. GEPA and SkillOpt answer a different question: whether a bounded candidate improves measured behavior. The durable architecture joins them without conflating them:

  1. capture a consented, scoped, content-addressed session;
  2. redact and classify prompts, tool arguments, and results before sharing;
  3. extract a candidate with exact evidence locators;
  4. deduplicate it against the current skill and other proposals;
  5. evaluate current versus candidate on a representative training set;
  6. select only on a separate held-out gate;
  7. publish a reviewed, versioned diff with rollback;
  8. project it through Skill Sync Workflow and prove routing/trigger behavior.

The capture/proposer may never edit the canonical skill directly. Recurrence is evidence that a behavior may be worth codifying, not evidence that the proposed wording is correct. Rejected candidates and failure traces remain negative training evidence. Source: frozen AMAP-ML/SkillClaw source and capture manifest, reviewed 2026-08-10

SkillClaw's present defaults also demonstrate why governance is separate from the collector: the proxy binds to all interfaces without an API key, records complete prompts and tool results, can upload them to shared storage, and can mutate several harness skill/config directories. A Kevin deployment remains held until loopback/auth, redaction and ACLs, immutable evidence, held-out evaluation, owner review, versioning, and rollback are proven.

Current Source Snapshot

As of 2026-07-01, gepa-ai/gepa is an MIT-licensed Python package with PyPI gepa@0.1.1, Python >=3.10,<3.15, GitHub HEAD 92dadfffbe98c8ecf508179a1cab09c1bb85cd32, tag v0.1.1, 5,464 stars, and 445 forks. The README now positions GEPA beyond prompts: textual system components can include prompts, code snippets, agent architectures, scheduling policies, vector graphics, MCP tool descriptions, DSPy programs, RAG prompts, TerminalBench agents, and LangChain/LangGraph pipelines via adapters. Source: PyPI gepa, 2026-07-01; Source: GitHub gepa-ai/gepa, 2026-07-01

As of 2026-08-11, microsoft/SkillOpt remains MIT-licensed with latest stable release v0.2.0 at e4ea6a6771e797ef820cdd8bfea64c57e0481065, published 2026-07-02. The observed main head was ba820b500f9da96685cf2780c7dc85ed4eb6563e, with 15,890 stars, 1,480 forks, 36 open issue/PR count, and an upstream push on 2026-08-09. Both the exact tag archive and the newer head archive are preserved; the head was inspected but not substituted for a release or installed. arXiv 2605.23904v2 was revised 2026-05-25. Source: GitHub API, git ls-remote, frozen archives, and arXiv v2, captured 2026-08-11

The supporting muratcankoylan/Agent-Skills-for-Context-Engineering evidence is MIT-licensed release v2.3.0 at 61f38ffc0ff3ae83adcf2fe011f3b751105add6d. At review time its current main head was 6dbe1a1d868eab51a3bc9011b0f55e2891513e40, with 17,688 stars and 1,457 forks. Kevin-Wiki preserved the exact v2.3.0 archive because that release owns the cited router methodology and results; mutable main discovery numbers are context, not quality proof.

The Akshay Pachaar bookmark is useful because it ties GEPA back into Hermes Agent (Nous Research) as one part of a compounding-agent stack: memory, self-evolving skills, profile-specific agents, model blending, and skill creation from examples. Its YouTube transcript is broader than GEPA, so keep this page focused on GEPA and route Hermes architecture details through Hermes Agent (Nous Research) and Agent Looping. Source: X/@akshay_pachaar, 2026-05-13; Source: YouTube transcript in enriched record, 2026-07-01


Timeline

  • 2026-08-11 | Replayed the SkillOpt source cluster against arXiv v2, exact v0.2.0, current main, its leak-abstention/evidence tests, and the exact Context Engineering Agent Skills v2.3.0 router benchmark. Corrected the X overclaims, separated description routing from body effectiveness, required per-skill effects, and implemented controlled-skill-optimization/v1 with five negative tests. No optimizer, plugin, provider, or self-modifier was installed. Source: X/@muratcan; frozen captures and local contract
  • 2026-08-10 | Added SkillClaw as the session-capture and candidate-proposal stage, explicitly separated it from GEPA/SkillOpt evaluation and promotion, and recorded the proxy/data/self-modification blockers that prevent global installation. Source: X/@aashatwt; frozen AMAP-ML/SkillClaw source
  • 2026-07-03 | Deep-reviewed the Muratcan Koylan SkillOpt bookmark and local diagram artifact, then checked current SkillOpt sources. The artifact's durable contribution is the optimizer-control shape: bounded skill edits, rejected side updates, held-out selection gates, and a text-space optimization analogy where edit budget functions like learning rate. Current source snapshot: microsoft/SkillOpt v0.2.0, MIT, HEAD e4ea6a6, 10,515 stars, 979 forks. Source: X/@muratcan, 2026-05-26; Source: GitHub microsoft/SkillOpt, 2026-07-03
  • 2026-07-01 | Deep-reviewed the Akshay Pachaar bookmark, YouTube crash-course transcript, arXiv v2, GitHub repo, and PyPI package. Current source snapshot: gepa@0.1.1, Python >=3.10,<3.15, MIT, tag v0.1.1, HEAD 92dadfffbe98c8ecf508179a1cab09c1bb85cd32, 5,464 stars, and 445 forks. Source: arXiv, GitHub, PyPI, 2026-07-01
  • 2026-05-27 | @kylejeong reports making browser skills up to 90% faster and cheaper via iterative AutoResearch, producing a /autobrowse command. 2,439 bookmarks. Source: X/@kylejeong, 2026-05-27
  • 2026-05-26 | SkillOpt (@muratcan, 2,283 likes, 5,050 bookmarks at capture): one of the first papers to treat markdown SKILL.md files as trainable parameters ("gradient descent for skill files"). Attacks skill optimization from the optimization end vs GEPA's governance end. Source: X/@muratcan, 2026-05-26
  • 2026-05-13 | Described in Akshay's Hermes Agent guide. Originally published as ICLR 2026 Oral paper by Nous Research. Source: X/@akshay_pachaar, 2026-05-13