LLM Wiki Agent Harness
Kevin-Wiki is Kevin's LLM wiki for an agent harness that acts as a second brain: Markdown is the durable source, agents are the maintainers, and every runtime gets the same identity, skills, memory, routing, and proof surfaces.
This page names the architecture Kevin is building. It is broader than a note vault and narrower than a full agent runtime like Hermes Harness (Nous Research). The wiki is the brain and harness preset that makes Cursor, Claude Code, Codex, Hermes, OpenClaw, and future agents behave like the same long-running system instead of unrelated chat sessions. Source: AGENTS.md; Cross-Agent Harness Portability
Boundary
The LLM wiki harness owns durable knowledge and operating policy. It does not own model inference, terminals, messaging delivery, or sandbox execution directly; those are delegated to the active runtime. Its job is to make every runtime load the same compiled brain and write back to the same source of truth.
| Layer | Local surface | Harness role |
|---|---|---|
| Identity spine | wiki/SOUL.md, wiki/USER.md, wiki/HEARTBEAT.md |
Compact boot context: who the agent is, who Kevin is, and what cadence applies. |
| Compiled knowledge | wiki/**/*.md, qmd, _index.md, backlinks |
Source-backed facts, decisions, concepts, project history, and routing context. |
| Raw evidence | raw/ plus cited URLs/assets |
Immutable source material that compiled pages draw from. |
| Procedural memory | skills/**/SKILL.md plus references/, templates/, scripts/, assets/ |
Repeatable workflows that change agent behavior. |
| Routing and authority | wiki/meta/, wiki/RESOLVER.md, wiki/SKILL-RESOLVER.md, capability maps |
Where knowledge belongs and what tool/skill should fire first. |
| Automation state | automations/, state.json, wiki/log.md |
Recurring work, freshness, receipts, and operation history. |
| Generated projections | skills/.runtime/all, wiki/skills/, registries, /api/system/graph |
Runtime and UI views generated from canonical skills and graph inputs. |
| Reader surface | ui/ and deployed site |
Human browsing, search, graph, and artifact review. |
Hermes Patterns Worth Emulating
Hermes provides a useful reference implementation because it separates the agent runtime into explicit surfaces instead of hiding everything inside chat memory.
| Hermes surface | What it does | Kevin-Wiki analogue |
|---|---|---|
| Progressive skill disclosure | skills_list shows names/descriptions; skill_view loads the full SKILL.md; support files load only when requested. Source: Hermes skills docs, 2026-07-01 |
Runtime skill index, generated skill registry, capability routing, and direct SKILL.md reads. |
| Skill packages | A skill is a directory with SKILL.md plus optional references/, templates/, scripts/, and assets/; writes are validated and can be approval-gated. Source: tools/skill_manager_tool.py, Hermes commit 9be292f |
Canonical skill source tree, runtime symlinks, skill-creator, and skills:check. |
| Memory split | Tiny curated MEMORY.md and USER.md, session FTS search, external memory providers, and write approval. Source: Hermes memory docs, 2026-07-01 |
Boot spine for compact identity, wiki/qmd for durable knowledge, logs/state for operations, and skills for procedures. |
| Context files | .hermes.md / AGENTS.md / CLAUDE.md / .cursorrules discovery, global SOUL.md, prompt-injection scan, and progressive subdirectory hints. Source: Hermes context-files docs, 2026-07-01 |
Root AGENTS.md, project agent-docs mesh, nearest context files, and config projections. |
| Toolsets | Named toolsets and platform presets, including constrained defaults for hostile surfaces such as webhooks. Source: toolsets.py, Hermes commit 9be292f |
Capability-routing map, harness-specific tool choices, and future per-surface authority profiles. |
| Cron tasks | Scheduled jobs run in fresh sessions, can attach skills, deliver to channels, run no-agent scripts, and chain context. Source: Hermes cron docs, 2026-07-01 | automations/, freshness checks, proof contract, and source-compile/radar jobs. |
| Gateway security | Allowlist-first messaging, DM pairing, approvals, blocklists, MCP env filtering, and context/memory injection scanning. Source: Hermes security docs, 2026-07-01 | Agent config doctors, process/git guards, capability routing, source provenance, and future approval surfaces for risky automation. |
Current LLM-Wiki conformance and extensions
The current Hermes llm-wiki skill remains a strong minimum contract, but it
is not the complete Kevin system. Use this comparison when source replay or UI
work risks collapsing the architecture back into “Markdown plus a graph.”
| Reference invariant | Kevin-Wiki implementation | Required proof |
|---|---|---|
| Immutable raw source body + hash | Content-addressed connector records, local media, repository/document captures, source inventory | Source/revision/hash/storage reference in a passing review |
| Resume through schema, index, recent log | SOUL/USER/HEARTBEAT, operations hub, status, qmd/index/backlinks, recent log | Boot/status doctor and current generated surfaces |
| Agent-owned entity/concept/comparison/query pages | Stable objects spanning knowledge, workflow, skill, capability, project, decision, and instruction owners | Object registry path and no duplicate owner |
| Page thresholds and split rules | Interrogate an existing owner first; create only a durable distinct object; retain broad sources as registries/corpora | Review verdict, comparison, rationale, and writeback digest |
| Confidence and contradictions stay visible | Signal confidence/novelty, counterexamples, incompatible source claims, held/blocked/watch states | Evidence locators and unresolved boundary preserved |
| Lint schema, links, orphans, staleness, duplicates, tags, sources | Wiki lint, source audits, skill/config/command doctors, brain doctor | Zero error rules plus explicit backlog warnings |
| Query or ingest updates the wiki | Named capture/research workflow interrogates stable owners and writes only behavior-changing knowledge | Passing review, promoted source row, index/embed/API/UI projection |
| Graph exposes relationships | Typed source → signal → object → workflow → proof graph | Graph helps retrieval/coverage; centrality never becomes truth or priority by itself |
The graph image in the saved post is therefore supported, but only as one
projection. It shows grpo, rl-post-training, index, atropos, and related
nodes forming a dense cluster while several source/schema/log nodes remain
disconnected. That topology can reveal coverage and possible missing links; it
cannot prove source completeness, correctness, freshness, causal importance,
or retrieval quality. Source: X/@Teknium graph and self-thread, replayed
2026-08-11
Design Rules
- The wiki and graph inputs are source; projections are build artifacts. If an index, API view, runtime symlink, or generated registry is wrong, fix the source or generator and rebuild. Source: Generated Surface Contract
- Memory layers must not collapse. Identity belongs in boot files, facts in wiki pages, procedures in skills, source evidence in
raw/, and current-task state in session context or logs. Source: Memory Routing Protocol - Procedures should be package-shaped. A mature skill should keep always-needed behavior in
SKILL.mdand move bulky examples, templates, scripts, or source excerpts into support directories. Source: Hermes skill authoring scan, 2026-07-01 - Authority is surface-specific. A wiki-ingest agent, cron job, code editor, messaging gateway, and webhook-triggered agent should not receive the same tool authority by default. Source: Hermes toolsets and webhook-safe tools, 2026-07-01
- Recurring work needs receipts. A scheduled job or skill update is not real until it leaves logs, outputs, state, generated surfaces, or eval receipts that the next agent can inspect. Source: Automation Proof Contract
- Corrections patch the failed procedure. If Kevin corrects style, workflow, or output shape, update the relevant skill or protocol, not only the user profile. Hermes' background review treats user frustration as a skill signal; this wiki should too. Source:
agent/background_review.py, Hermes commit9be292f - Workflows are the execution entry point. The brain should route a job to a named workflow that declares its knowledge, patterns, skills, tools, authority, proof, and writeback contract. Source: User's Straight Fats architecture notes, 2026-07-11 through 2026-07-14
- Selected examples become fixtures. A user-selected X post or artifact is integrated only when a workflow owns its useful pattern, can evaluate it, and records adoption or rejection. Presence in an ingest ledger alone is not sufficient. Source: User clarification, 2026-07-14
- Citations verify the source text, not only the URL. Quote-bearing or decisive claims need evidence that the cited body contains the claimed text; inaccessible or changed pages remain unresolved. Source: Hermes grounded-citations release contract, v0.20.0
- Events are authenticated and replay-safe. Runtime lifecycle events may project through signed webhooks or A2A, but receivers verify signatures, dedupe event IDs, restrict authority, and write canonical outcomes back through an owned workflow. Source: Hermes signed webhooks and A2A v1.0, v0.20.0
- Recovery preserves evidence. Truncated tool output spills to a readable artifact, writes are verified, already-applied patches become idempotent success, and near-miss searches remain diagnostics—not silent guesses. Source: Hermes tool self-recovery wave, v0.20.0
- Compaction cannot erase active instructions. Protect recent user turns, mark pruned skills deterministically, redact at every compaction boundary, and soft-defer when state locks rather than inventing a partial summary. Source: Hermes compression/security changes, v0.20.0
- Approval fatigue fails closed. Approval history may suggest narrow allowlists, but repeated denials trigger a circuit breaker and never auto-expand authority. Source: Hermes smart-approval changes, v0.20.0
How It Differs From Hermes
Hermes is the runtime: it owns the agent loop, messaging gateway, terminal backends, live tools, session DB, cron scheduler, and profile-scoped ~/.hermes state. Kevin-Wiki is the brain and harness preset: it owns durable meaning, source evidence, cross-harness skill projection, and the operating model that multiple runtimes inherit. Source: Hermes architecture docs, 2026-07-01; Cross-Agent Harness Portability
That difference matters. Hermes keeps memory compact because it enters every system prompt. Kevin-Wiki can be large because the active agent searches and loads only relevant pages, objects, workflows, and skills. Hermes writes memories and skills into one profile. Kevin-Wiki writes canonical owners that many harnesses query through the same graph and routes. Source: Hermes memory docs, 2026-07-01; Wiki Brain Operating Model
External Confirmation
Matt Pocock described the same pattern publicly on 2026-07-01: a personal, entirely agent-managed Karpathy-style wiki ingesting X, Discord, and Gmail every few hours, then serving as the knowledge base/environment for future loops. He later framed the payoff as course planning with an agent that understands the whole business context, including asking it about student feedback, missing talking points, previous cohort problems, and current learner demand. Source: X/@mattpocockuk, 2026-07-01
For Kevin-Wiki, that is a peer signal rather than a new architecture. It reinforces three existing rules: ingest live sources into durable pages, compile business/personal context into searchable articles instead of chat memory, and turn repeated work into loops, skills, or automations. Source: Capture Ingest Protocol; No One-Off Work; The Brain-Agent Loop
Implementation Implications
| Need | Current implementation | Next pressure |
|---|---|---|
| Fast boot | SOUL/USER/HEARTBEAT plus operations hub | Keep boot files compact and push expansions into subpages. |
| Reliable retrieval | qmd, _index.md, backlinks, brain graph and APIs |
Add quote-to-source verification and retrieval receipts where decisive claims warrant them. |
| Cross-harness skills | skills/.runtime/all and ~/.agents/skills bridge |
Treat external harness-managed skills as read-only inputs unless promoted. |
| Repeated capture | capture ingest protocol, x-bookmark audits, log | Use hash/drift checks for raw URL captures when practical. |
| Scheduled work | automations/, freshness, proof contract |
Borrow Hermes' no-agent watchdog pattern for deterministic alerts. |
| Tool authority | capability routing map and harness projections | Add per-surface profiles for webhook, A2A, cron, messaging, repo edit, and browser work; signed transport never implies execution authority. |
| Human review | log, doctors, generated audit queues | Add approval queues and repeated-denial circuit breakers where automation can mutate identity, skills, or public output. |
| Runtime recovery | resumable workflows, artifacts, idempotent source state | Standardize overflow artifacts, post-write verification, already-applied success, and near-miss diagnostics. |
| Context compression | boot spine plus runtime-specific compaction | Require recent-user-tail survival, deterministic pruned-skill markers, redaction, and lock-safe deferral. |
Failure Modes
| Failure | Result | Fix |
|---|---|---|
| Treating the wiki like RAG storage | Agents re-derive knowledge and miss routing decisions. | Compile facts into canonical owners and update qmd and graph surfaces. |
| Copying skills per runtime | Cursor, Codex, Hermes, and Claude drift. | Keep one skill source tree and project through symlinks or configured external dirs. |
| Saving everything into boot memory | Startup context bloats and becomes stale. | Keep boot memory compact; move detail to wiki pages, skills, or qmd-searchable logs. |
| Cron without proof | Scheduled work is invisible to the next session. | Require output artifacts, state updates, log entries, and promotion/no-op decisions. |
| Hostile input receives full tools | Webhooks, public chats, or untrusted repos can steer powerful tools. | Use constrained toolsets, approvals, context scans, and allowlists. |
| User corrections stay conversational | The same failure repeats later. | Patch the responsible skill/protocol and regenerate generated surfaces. |
Timeline
- 2026-08-11 | Re-interrogated the architecture against current Hermes LLM-Wiki v2.1.0 and Hermes Agent v0.20.0. Added the source/object/workflow/proof conformance table, graph-evidence boundary, quote verification, signed replay-safe events, tool-recovery receipts, compaction invariants, and approval-denial circuit breaking. Source: X/@Teknium;
NousResearch/hermes-agent@c0106e50; releasev2026.8.3 - 2026-08-10 | Made the graph-backed brain and named workflows the complete execution and context model. Source: User request, 2026-08-10
- 2026-07-01 | Added Matt Pocock's agent-managed Karpathy-style wiki post as external confirmation for this architecture: live source ingests become the durable context environment for future agent loops and domain-specific planning. Source: X/@mattpocockuk, 2026-07-01
- 2026-07-01 | Page created after Kevin asked to study Hermes Agent's skills/docs/code and emulate the important pieces in "my LLM wiki for an agent harness that acts as a second brain." Reviewed Hermes docs,
skills/,toolsets.py, skill manager, prompt builder, memory/cron/security/context docs, and main commit9be292f1e678437644396b47b3410b433ba3433f; promoted the durable architecture into Kevin-Wiki. Source: User request, 2026-07-01; Hermes docs and repo, 2026-07-01