Controlled Skill Evolution - GEPA and SkillOpt
Skills may improve from execution evidence, but only through a controlled candidate-selection-promotion loop. GEPA explores textual variants through Genetic-Pareto search; SkillOpt trains one compact skill through bounded patches and strict held-out selection. Neither licenses unconditional self-rewrite. [Sources: GEPA; SkillOpt]
The Problem It Solves
In-agent skill learning has a known weakness: agents tend toward self-congratulation. They almost always rate their own performance highly even when they failed. The same system that auto-generates skills can overwrite manual customizations with worse versions. Community feedback confirms this is a real and persistent failure mode. Source: compiled from X/@akshay_pachaar article, 2026-05-13
How It Works
- Read current skill from the agent's skill directory
- Generate evaluation dataset (synthetic via Claude Opus, real session history from SQLite, or hand-curated golden sets)
- Run GEPA optimizer: read execution traces -> understand failure points -> generate candidate variants through evolutionary search
- Evaluate candidates using LLM-as-judge scoring with rubrics (not binary pass/fail)
- Apply constraint gates: full test suite passes 100%, skills stay under 15KB, caching compatibility preserved, semantic purpose doesn't drift
- Best variant goes out as a PR - never a direct commit
Cost: $2-10 per optimization run. No GPU required - everything through API calls.
Relevance to Kevin's System
Kevin's skill-maintenance automation already runs periodic skill updates from session analysis. GEPA offers a more rigorous alternative: instead of the agent deciding "I should improve this skill," execution traces objectively show where skills failed, and evolutionary search finds better variants with constraint preservation.
The PR-based output (never direct commit) aligns with Kevin's code review conventions. The self-congratulation problem is exactly why Kevin uses Security and Review Skills - a second model catches what the first model thinks is fine.
GEPA is positioned as a middle ground: more rigorous than in-agent self-improvement, cheaper than full RL fine-tuning (GRPO). The paper's current arXiv v2 says GEPA reflects over trajectories, proposes prompt updates, combines complementary lessons from the Pareto frontier, outperforms GRPO by 6% on average and up to 20%, and uses up to 35x fewer rollouts. Source: arXiv:2507.19457v2, 2026-02-14
SkillOpt sharpens the same skill-evolution problem from the optimizer side. Its diagram treats a skill document as the trainable state of a frozen agent: rollouts produce trajectories, a separate optimizer proposes bounded add/delete/replace edits, a held-out gate accepts only strict validation improvements, and rejected side updates become negative feedback. This is closer to text-space gradient descent than GEPA's broader Genetic-Pareto prompt evolution, but both reinforce the same operational rule for Kevin's system: optimize skills only through trace evidence and validation gates, never through unconditional self-rewrite. Source: X/@muratcan and local image artifact, replayed 2026-08-11; Source: SkillOpt project page
Controlled Skill Optimization Contract
Kevin-Wiki implements the shared lesson as a stricter local contract:
The machine-readable form is skillOptimization in
evals/agent-self-improvement/suite.json; eleven contract tests now cover it
and the controlled-evaluator contract together. The target model, execution
harness, evaluator, and split revisions stay frozen while one skill candidate
is compared. Training creates candidates; validation selects; final test only
reports. Ties reject, holdout leakage abstains, and the canonical skill cannot
change before the gate. Fast patches may not overwrite protected slow policy.
Open-ended writing, design, and strategy updates remain held without a reliable
verifier or reviewed evidence. Source: local contract and tests, implemented
2026-08-11
Two independently failing surfaces
A skill has at least two measured surfaces:
| Surface | What the agent sees | Required evidence |
|---|---|---|
| Router description | Frontmatter names/descriptions only | top-1/top-k routing, false triggers, near-misses, and per-skill confusion |
| Activated body | Full selected SKILL.md and resources |
task/state outcome, regressions, cost, and negative transfer |
| End-to-end | Normal runtime discovery and activation | the intended description selects the intended body and improves the task |
The v2.3.0 Context Engineering Agent Skills benchmark proves why those lanes
cannot be merged. Its router sweep deliberately used settingSources: [], so
the model saw descriptions but no bodies. Targeted description changes moved
context-fundamentals by +23.4 percentage points, project-development by
+25, and tool-design by +7.8, while whole-model top-1 deltas were only +2.5,
+2.7, +3.9, and -2.0 points across the four reported models. The repository
explicitly labels body-loaded effectiveness as a later stage that had not yet
run. Aggregate accuracy would have hidden the important skill-level effects;
router accuracy would have overstated what had actually been proven. Source: v2.3.0 release,
README, changelog, router runner, and published 600-run report; captured
2026-08-11
What SkillOpt Actually Shows
The arXiv v2 protocol uses four epochs, a default rollout batch of 40, reflection minibatches of 8, and a textual edit budget of 4 with cosine decay to 2. Those are experiment defaults, not universal constants. The learning-rate sweep has mixed winners across SearchQA, SpreadsheetBench, and LiveMath; it does not establish a general “4–8 edits is the sweet spot.” Removing the budget made the tested scores worse, not universally collapse. The protected slow/meta ablation's largest drop was SpreadsheetBench 77.5 to 55.0 (-22.5 points); that is a benchmark-specific result, not a universal 22-point law. Source: SkillOpt arXiv v2, Ablations
The paper's learned GPT-5.5 artifacts were 379–1,995 tokens with a roughly 920-token median and only 1–4 accepted updates. That is evidence for compact, selective artifacts, not a target length or edit count that every skill must copy. The clearest portability result is SpreadsheetBench: Codex-trained to Claude Code improved 22.1 to 81.8 (+59.7), while Claude-Code-trained to Codex improved 27.5 to 71.1 (+43.6). LiveMath cross-harness gains were smaller (+1.6 and +12.8), so every destination still needs its own transfer receipt. The authors report best or tied-best results on 52/52 evaluated cells; that claim is bounded to their models, benchmarks, scorers, splits, and harness adapters.
The method depends on scored trajectories and a disjoint selection split. Open-ended subjective work needs stronger human or model evaluation; training adds rollout and optimizer cost; one compact skill can be insufficient for a heterogeneous domain; and transfer still requires careful held-out testing. The deployed skill has no optimizer call, but “zero inference-time cost” means no extra model call—not zero prompt tokens. Source: SkillOpt arXiv v2, Experiments, Limitations, and Conclusion
Adoption Boundary
| Borrow | Do not infer or adopt |
|---|---|
| bounded, provenance-bearing patches | autonomous canonical rewrites after a single success |
| strict disjoint held-out improvement; ties reject | aggregate-only promotion or leaked holdouts with a warning |
| rejected-edit ledger and protected slow state | treating all rejected edits as bad when the gate itself leaked |
| separate router, body, and end-to-end lanes | calling description accuracy skill effectiveness |
| per-skill effects, regressions, and transfer checks | universal edit-count, token-length, or cross-harness guarantees |
| exact pinned source as research/reference input | installing SkillOpt-Sleep globally or authorizing provider spend |
The current SkillOpt main branch adds useful holdout-leak abstention tests and
append-only evidence machinery beyond the stable v0.2.0 tag. It also contains
an optional gate bonus based on “semantic density” of words such as MUST,
ALWAYS, and CRITICAL. Kevin-Wiki does not adopt that bonus: instruction-word
density is a gameable textual proxy and cannot substitute for routing or task
outcomes. This is a local inference from the inspected current source, not a
claim by the paper. The repository remains a pinned research reference rather
than a globally installed self-modifier.
Session Capture Is Not Skill Promotion
SkillClaw adds a useful upstream stage: it can intercept model/tool trajectories, collect repeated strategies across sessions, summarize candidate skills, and optionally distribute them through a shared registry. GEPA and SkillOpt answer a different question: whether a bounded candidate improves measured behavior. The durable architecture joins them without conflating them:
- capture a consented, scoped, content-addressed session;
- redact and classify prompts, tool arguments, and results before sharing;
- extract a candidate with exact evidence locators;
- deduplicate it against the current skill and other proposals;
- evaluate current versus candidate on a representative training set;
- select only on a separate held-out gate;
- publish a reviewed, versioned diff with rollback;
- project it through Skill Sync Workflow and prove routing/trigger behavior.
The capture/proposer may never edit the canonical skill directly. Recurrence is evidence that a behavior may be worth codifying, not evidence that the proposed wording is correct. Rejected candidates and failure traces remain negative training evidence. Source: frozen AMAP-ML/SkillClaw source and capture manifest, reviewed 2026-08-10
SkillClaw's present defaults also demonstrate why governance is separate from the collector: the proxy binds to all interfaces without an API key, records complete prompts and tool results, can upload them to shared storage, and can mutate several harness skill/config directories. A Kevin deployment remains held until loopback/auth, redaction and ACLs, immutable evidence, held-out evaluation, owner review, versioning, and rollback are proven.
Current Source Snapshot
As of 2026-07-01, gepa-ai/gepa is an MIT-licensed Python package with PyPI gepa@0.1.1, Python >=3.10,<3.15, GitHub HEAD 92dadfffbe98c8ecf508179a1cab09c1bb85cd32, tag v0.1.1, 5,464 stars, and 445 forks. The README now positions GEPA beyond prompts: textual system components can include prompts, code snippets, agent architectures, scheduling policies, vector graphics, MCP tool descriptions, DSPy programs, RAG prompts, TerminalBench agents, and LangChain/LangGraph pipelines via adapters. Source: PyPI gepa, 2026-07-01; Source: GitHub gepa-ai/gepa, 2026-07-01
As of 2026-08-11, microsoft/SkillOpt remains MIT-licensed with latest stable
release v0.2.0 at e4ea6a6771e797ef820cdd8bfea64c57e0481065, published
2026-07-02. The observed main head was
ba820b500f9da96685cf2780c7dc85ed4eb6563e, with 15,890 stars, 1,480 forks,
36 open issue/PR count, and an upstream push on 2026-08-09. Both the exact tag
archive and the newer head archive are preserved; the head was inspected but
not substituted for a release or installed. arXiv 2605.23904v2 was revised
2026-05-25. Source: GitHub API, git ls-remote, frozen archives, and
arXiv v2, captured 2026-08-11
The supporting muratcankoylan/Agent-Skills-for-Context-Engineering evidence
is MIT-licensed release v2.3.0 at
61f38ffc0ff3ae83adcf2fe011f3b751105add6d. At review time its current main
head was 6dbe1a1d868eab51a3bc9011b0f55e2891513e40, with 17,688 stars and 1,457
forks. Kevin-Wiki preserved the exact v2.3.0 archive because that release owns
the cited router methodology and results; mutable main discovery numbers are
context, not quality proof.
The Akshay Pachaar bookmark is useful because it ties GEPA back into Hermes Agent (Nous Research) as one part of a compounding-agent stack: memory, self-evolving skills, profile-specific agents, model blending, and skill creation from examples. Its YouTube transcript is broader than GEPA, so keep this page focused on GEPA and route Hermes architecture details through Hermes Agent (Nous Research) and Agent Looping. Source: X/@akshay_pachaar, 2026-05-13; Source: YouTube transcript in enriched record, 2026-07-01
Timeline
- 2026-08-11 | Replayed the SkillOpt source cluster against arXiv v2, exact
v0.2.0, current
main, its leak-abstention/evidence tests, and the exact Context Engineering Agent Skills v2.3.0 router benchmark. Corrected the X overclaims, separated description routing from body effectiveness, required per-skill effects, and implementedcontrolled-skill-optimization/v1with five negative tests. No optimizer, plugin, provider, or self-modifier was installed. Source: X/@muratcan; frozen captures and local contract - 2026-08-10 | Added SkillClaw as the session-capture and candidate-proposal stage, explicitly separated it from GEPA/SkillOpt evaluation and promotion, and recorded the proxy/data/self-modification blockers that prevent global installation. Source: X/@aashatwt; frozen AMAP-ML/SkillClaw source
- 2026-07-03 | Deep-reviewed the Muratcan Koylan SkillOpt bookmark and local diagram artifact, then checked current SkillOpt sources. The artifact's durable contribution is the optimizer-control shape: bounded skill edits, rejected side updates, held-out selection gates, and a text-space optimization analogy where edit budget functions like learning rate. Current source snapshot:
microsoft/SkillOptv0.2.0, MIT, HEADe4ea6a6, 10,515 stars, 979 forks. Source: X/@muratcan, 2026-05-26; Source: GitHubmicrosoft/SkillOpt, 2026-07-03 - 2026-07-01 | Deep-reviewed the Akshay Pachaar bookmark, YouTube crash-course transcript, arXiv v2, GitHub repo, and PyPI package. Current source snapshot:
gepa@0.1.1, Python>=3.10,<3.15, MIT, tagv0.1.1, HEAD92dadfffbe98c8ecf508179a1cab09c1bb85cd32, 5,464 stars, and 445 forks. Source: arXiv, GitHub, PyPI, 2026-07-01 - 2026-05-27 | @kylejeong reports making browser skills up to 90% faster and cheaper via iterative AutoResearch, producing a /autobrowse command. 2,439 bookmarks. Source: X/@kylejeong, 2026-05-27
- 2026-05-26 | SkillOpt (@muratcan, 2,283 likes, 5,050 bookmarks at capture): one of the first papers to treat markdown SKILL.md files as trainable parameters ("gradient descent for skill files"). Attacks skill optimization from the optimization end vs GEPA's governance end. Source: X/@muratcan, 2026-05-26
- 2026-05-13 | Described in Akshay's Hermes Agent guide. Originally published as ICLR 2026 Oral paper by Nous Research. Source: X/@akshay_pachaar, 2026-05-13