Capability Routing Map

Skills are procedures. Tools are capabilities. Duplicate skill/tool pages are valid when they document different layers: the skill tells the agent how to run the workflow; the tool page tells the agent what the capability is, when it wins, and what replaces it. Source: User, 2026-06-17

This page is harness-agnostic. Cursor, Codex, Claude Code, OpenClaw, Hermes, and future agents should treat this page plus Skill Resolver as the canonical tool-routing brain. Files such as config/cursor/rules/tool-hierarchy.mdc are compact projections for a specific loader, not independent sources of truth.

Decision Rule

When a task overlaps multiple tools, route by capability family first, then by task shape, then by interface cost.

Discovery happens one layer earlier in Agent Capability Registry: what exists, how it is invoked, what it requires, and who owns it. This map decides what to use first.

Coverage happens alongside discovery in Ecosystem Capability Coverage: whether an ecosystem is present in skills.sh/MCP/CLI sources, promoted into wiki pages, routed in the resolver, projected into active runtime indexes, and backed by verification. Run that audit before large page consolidation or runtime-projection edits so the wiki does not over-promote whichever vendor family happened to be documented first.

  1. Family - search, browser/computer use, service/API access, docs lookup, model choice, QA, content, data movement, or memory.
  2. Task shape - deterministic vs exploratory, current vs compiled, repeated vs one-off, shell-agent vs IDE-agent, human-facing vs internal.
  3. Interface cost - prefer the narrowest reliable interface: local wiki/search or CLI before MCP; MCP before raw HTTP; browser before computer-use screenshots only when interaction is unavoidable.

When to open this map

Open this page when the tool question is not a simple service lookup. Common triggers:

  • "Should I use Playwright, agent-browser, or Chrome DevTools?"
  • "Should I search the web, qmd, or last30days?"
  • "Should this be a skill, tool page, concept page, or resolver row?"
  • "Should I use a CLI, MCP, browser, or raw API?"
  • "Which agent runtime/model path fits this constraint?"

If the task is already covered by a specific skill with no ambiguity, read that skill instead of over-routing. This map is for capability boundaries.

Cost Ladder

Prefer the lowest-cost interface that preserves correctness:

  1. local files / rg / qmd
  2. deterministic scripts or CLI
  3. official docs or local docs mirrors
  4. MCP tools
  5. browser automation
  6. raw computer-use / screenshot driving
  7. custom one-off code

Drop lower only when the current rung cannot answer the task.

Skill vs Tool Layer Rule

The wiki no longer mirrors every executable skill as wiki/skills/foo.md. Executable procedures live in skills/**/SKILL.md, inventory lives in Skill Registry, readable skill-family synthesis lives in wiki/concepts/*-skills.md, and tool pages own product/capability facts.

Layer Owns Example
Executable skill Trigger phrases, preconditions, step-by-step workflow, required reads, verification, output shape skills/engineering/agent-browser/SKILL.md: how to run browser automation in an agent loop
Skill family page Cross-skill routing, family tradeoffs, cleanup rules, notable procedures Browser Testing Skills
Tool page Product/capability facts, tradeoffs, alternatives, timelines, related tools [[agent-browser]] tool: CLI/MCP surfaces, Lightpanda/Chrome engines, React/perf strengths
Concept page General principle that cuts across tools [[computer-use-patterns]]: browser-use spectrum and persistent playbook pattern
Resolver row Which one fires first for a concrete user request [[SKILL-RESOLVER]] Browser Automation table

Do not recreate a skill article only because a skill exists. Promote durable ideas into the right family page, tool page, or concept page; preserve old slugs with redirects.

Core Capability Families

Family First route Tool hierarchy Drop rank when
Kevin/wiki knowledge The Brain-Agent Loop qmd search for exact names → qmd query for synthesis → read pages + backlinks → web only for gaps qmd lacks current/external facts
Source-to-skill compilation absorb-sources → skill-creator freeze and interrogate the source → confirm the signal is a recurring procedure or bounded knowledge base → extend an existing skill before creating one → use Hermes /learn or equivalent to compile a staged candidate → inspect diff → execute its verification → prove description routing and body behavior separately → promote the exact passing digest An ordinary fact/preference/design/tool signal is being turned into skill sprawl; source text gains instruction authority; a ## Verification section or launch claim is treated as an executed test; runtime write access bypasses canonical review; approval is assumed on; no receipt or rollback exists
Agent runtime memory Agent Memory System Architecture + Hindsight Memory Provider immutable source + Kevin-Wiki/qmd canonical owner → Hermes built-in session memory when sufficient → Hindsight recall only after pinned install, declared model/data path, per-user/repo bank map, doctor, leakage, fixed-corpus usefulness, and fresh-target export/replay proof → Hindsight reflect only for explicit synthesis with citations and cost/latency receipt → Honcho only for separately scoped peer/profile modeling A selected provider is being called active without machine proof; a loopback API is called fully local while inference leaves the machine; one shared bank lacks strict tags/provenance/leakage tests; recall and reflect costs are conflated; raw source or canonical rules would exist only inside runtime memory; benchmark protocols are unmatched
Query-embedding adaptation Retrieval Quality exact/alias/chunk/freshness/ACL fixes first → freeze a labeled qmd failure benchmark and identity baseline → trial a query-side linear adapter with leakage-free splits and mixed negatives → compare exact, base semantic, adapted, and hybrid lanes on per-family precision/recall/MRR/nDCG plus answer evidence, latency, ACL, and provenance → promote only the reversible passing digest No systematic query/corpus mismatch is proven; a template/hash demo, train score, or corpus average replaces held-out real-query evidence; exact-name/protected-query regressions or leakage are hidden; corpus/embedder/labels changed without re-evaluation
Relationship-heavy Postgres retrieval Retrieval Quality ordinary SQL/joins first → pgGraph only for repeated bounded traversal over verified relationships → Polygres/pgContext only when graph + lexical/vector hybrid retrieval is a measured product need Markdown/wiki search, simple joins, or any private/multi-tenant topology whose RLS/tenant isolation has not been proven
Rendered-layout retrieval Retrieval Quality + PixelRAG establish exact/text/DOM baseline → add screenshot-tile retrieval when tables, charts, diagrams, dashboards, PDF layout, or rendered state carry the answer → preserve source revision + render-state manifest + tile/region + link metadata → compare or fuse lanes under one reader/query/budget receipt Text/DOM already answers reliably; auth/actions/timing/locale/viewport are unrecorded; region grounding, cross-modal disagreement, privacy, rights, moderation, reader capability, storage/latency, or bounded-corpus quality proof is missing
Repo topology / code graph rg + Graphify + graphify-sidecar skill exact/source-known question → rg and raw reads → repeated project-specific F#/spec/glossary/outline/history question → a reviewed discover.sh-style wrapper with transparent freshness and false-positive limits → unfamiliar topology/path/explain/affected → npm run graphify:sidecar -- status then build/query → Code Review Graph only for a project-owned Tree-sitter/MCP or incremental-review need A custom wrapper merely renames generic search; the question is compiled wiki truth; an index is stale/noisy; freshness, context reduction, maintenance, and failure behavior were not compared; CRG config/hooks lack approval
Pull-request and diff review Security and Review Skills freeze base + Standards and Spec sources → inventory and hash the whole change surface → optionally run pinned meat on a large repetitive agent-written diff for a conceptual reading view → always compare that abridgment with the original path inventory and exact diff → for large/heterogeneous diffs or explicit full coverage, add OpenCodeReview delegation preview/rules when an audited ocr is available → account for every (path,status) as reviewed or skipped with reason → inspect, falsify, report → improve reviewer policy only through separate evidence gates A compressed reading view is presented as full coverage; meat drops an independent route, security, authority, or failure change behind ...; sensitive diffs leave the approved provider boundary; preview exclusions, untracked files, recreated paths, or skipped rows disappear; OCR benchmark claims are treated as independent proof; global install, provider credentials, comment publication, or fixes inherit read-only review authority
CAD, robot description, and fabrication artifacts CAD and Fabrication Workflow freeze units, tolerances, assumptions, deliverable, and authority → route only the needed cad / cad-viewer / step-parts / dxf / urdf / srdf / sdf / implicit-cad / gcode / sendcutsend / bambu-labs skills → preserve editable source and STEP-first primary geometry where applicable → deterministic validation → visual inspection → digital handoff or dry-run physical plan A render is called manufacturing proof; implicit CAD silently replaces STEP-first source; robot frames/limits/inertials/collisions are guessed; slicer or material profile is invented; download, vendor upload, purchase, machine motion, or print start lacks exact authority
Self-hosted visual CMS existing project-native CMS or static source first retain Kevin-Wiki's custom viewer → Instatic only when the product needs one typed editable page tree shared by visual editing, agent tools, content/data workspaces, and clean static publishing → pin release, prove import/export and backup/restore, database parity, capability gates, plugin sandbox, E2E, accessibility, performance, and rollback A pre-1.0 visual CMS replaces a working source-owned system from screenshots or stars; hostile multi-user deployment, plugin permissions, auth, backups, migrations, generated-output ownership, or source replay are unproved
Binary-heavy version control Git first; Lore Version Control candidate ordinary code/docs/wiki → Git; existing large-binary team → incumbent Git LFS or Perforce baseline; binary-dominant game/media/3D repository with measured clone, transfer, dedup, sparse-hydration, or offline pain → pinned Lore project pilot with representative corpus, concurrent locking/conflict, auth/ACL/isolation, backup/restore, upgrade, performance/cost, export, and rollback proof The repository is mostly text; a launch post or star count is the only need; Lore's pre-1.0/API/protocol, advisory-lock, UEFN-compression, server topology, security, recovery, or migration boundaries are unproved; do not install globally or replace Git by default
Current internet/social research last30days skill last30days → Agent Operations Skills channel router → Exa Agent for async structured list/enrichment research → WebSearch / source-specific browsing Need exact current docs, a primary source, or low-latency fetch
macOS app capability discovery workflow requirement + npm run harvest:mac-apps name the missing job/profile → snapshot and diff Awesome-Mac candidates → filter by category/labels → dedupe against installed owners → verify the selected app's official source, license, maintenance, signing/notarization, OS/hardware, price, permissions/data, package source, and uninstall path → bounded trial → doctor → workflow bind A popular catalog row, open/free/native label, sponsor placement, star count, or catalog license is being treated as app trust or install authority; no named workflow needs the app
Mixed-risk external catalog discovery named legitimate capability gap + npm run harvest:fmhy retain and diff the content-addressed file manifest → keep candidate projection disabled → open one relevant review-first source slice → verify shortlisted targets at official sources → check rights/license/maintenance/security/overlap → promote through the normal owner and trigger proof The request seeks unauthorized access; the catalog is being bulk-imported or republished; a source star/safety label becomes local trust; SafeGuard or generated bookmarks are installed globally; no bounded workflow need exists
Design-tool discovery and DOM-to-design capture Design Engineering Resources + Designtools.fyi name the product/design job → use Design Engineer Tools as the broad shelf and designtools.fyi as the role-ranked shortlist → open each selected tool's official repository, license, release, source, tests, data boundary, and runtime → use Fieldwork — Visual Design IDE for the composed semantic capture/artboard/version/proof surface → DialKit for animation inventory/timeline/control → MIT Tweakpane only for narrow parameter/monitor panes → Agentation only as a PolyForm-Shield-guarded internal/reference adapter, never as competing Fieldwork code without separate permission → require capture identity, component/source resolution method and confidence, SSR/non-React fallbacks, isolated non-mutating state, and independent acceptance proof A catalog score, “open source” label, stars, screenshot layer tree, React-only happy path, or self-resolved fixer queue is being treated as adoption proof; component/source identity is invented; the artboard mutates product source before approval; Agentation's competition restriction is ignored; Tweakpane is being promoted into the product state model
Scroll-scrubbed cinematic world scroll-world + Scroll World user asks for a fly-through brand world, isometric diorama journey, or scroll cinematic → distinguish pre-rendered scroll-controlled video from runtime 3D and ordinary DOM choreography → freeze story, brand, camera grammar, desktop/mobile scope, source-asset rights, provider/model, price and spend approval → qualify first/last-frame conditioning before the batch → generate frame- and velocity-compatible legs/connectors → encode seekable desktop plus approved native 9:16 mobile chain → preserve static/reduced-motion document → prove seams in both scroll directions, mobile delivery/decoding, keyboard/readability, teardown, cost and independent visual acceptance Interactive geometry/camera/state is required; GSAP/Motion can express the DOM behavior directly; the deliverable is a standalone video; provider auth, data egress, asset rights, price, or spend approval is missing; a still/PSNR alone substitutes for motion review; landscape crop is called mobile; a failed connector silently becomes a direct crossfade; repository prose is treated as target-app proof
Procedural Three.js environment threejs-towers / threejs-landscape / threejs-weather + Web3D Library Ranking name the subject, stage and atmosphere jobs separately → freeze owned/reference assets and licenses → use only the needed MIT skill modules → keep geometry, landscape and weather behind explicit scene-state seams → expose inspectable time/weather/timeline controls → prove winding/caps, shared-material draw calls, disposal, height-function agreement, frustum density, interruption, mobile input, audio/source rights, semantic fallback, reduced motion and target-device performance Static image/video or ordinary DOM is sufficient; one copied showcase monolith replaces the three modules; a public repository is assumed licensed; generated textures/audio lack reuse rights; performance is inferred from triangle count alone; no semantic/static fallback exists
Runnable AI app and agent-pattern discovery retained source catalog + named product or workflow gap Awesome LLM Apps as a high-signal example feed → choose one exact subdirectory by job → preserve the pinned revision and top-level plus nested licenses → inspect dependencies, providers, models, secrets, data, media, side effects, tests, evals, deployment, cost, and overlap → run in a disposable environment → compile the transferable pattern into the existing owner or a distinct audited skill Repository popularity or a top-level Apache license is being treated as blanket approval; the whole catalog is cloned into active runtime; examples, generated output, provider credentials, or nested assets are assumed production-ready; no exact app, gap, evaluator, or owner is named
Structured web extraction / crawl Firecrawl known readable URL: direct/Jina/Defuddle - Web Content Extraction → Firecrawl scrape for rendered/schema output → search for unknown URL → map for URL discovery → bounded crawl/batch → interact only for necessary dynamic actions → Browser Testing Skills for unconstrained authenticated work One direct fetch is sufficient; target access/terms, hosted-data handling, credit/page budget, or self-host auth is unresolved
Recorded media / meetings Agent Operations Skills Native captions or platform export first → direct pinned yt-dlp/ffmpeg for reproducible authorized capture → local OpenAI Whisper for offline file ASR when native text is absent → ReClip - Self-Hosted Video Downloader for a human-operated loopback batch/quality UI → Recordly for consented screen/product capture → audited Claude Video watch candidate only after its supply-chain hold clears; evaluate Meetily for local live meeting capture/transcription/summaries with explicit consent, storage, model, and provider-egress settings The source is text-native; download/reuse rights are unresolved; language/model/VRAM/accuracy is unmeasured; diarization or live meeting semantics are being inferred from Whisper; ReClip is exposed beyond loopback; cloud transcription was not approved; unaudited Recordly extensions are enabled; “local” is being inferred while an external summary provider is configured
Realtime voice and wearable multimodal agents Realtime Voice Agent Workflow freeze user job, authority, processors, client/transport, VAD, STT/native audio, realtime model, action model/gateway, tools, TTS/voice, sensor semantics, hardware/network, identity, and data policy → TypeScript/Next managed-provider surface: Vercel AI SDK + AI Gateway → OpenAI-centered behavior: OpenAI Agents SDK → Pipecat for a modular Python pipeline spanning many STT/LLM/TTS/S2S providers, WebRTC/WebSocket/telephony transports, and multi-agent handoffs → Hugging Face Speech to Speech when self-hosted/local components or its OpenAI-Realtime-compatible client are the measured need → wearable/first-person architecture and counterexample reference: VisionClaw → draw separate perception and action loops → bind short-lived scoped session identity → explicit current/pinned visual reference → human-only confirmation → durable late work → deterministic protocol/state tests → representative latency/quality/interruption/tool/privacy/security/cost matrix → low-authority canary → receipt-bearing promotion The job is recorded-file transcription; a polished demo or median replaces p95 and target-corpus proof; README, platform, tag, and moving main are mixed; “open” or “local” ignores license, terms, models, or providers; a model/transport approves its own always_ask event; room/media grants, user binding, consent/recording/retention/deletion, barge-in/stale output, job restart/single delivery, liveness/degraded state, accessibility, capacity, cost, or rollback is unproved
External docs / APIs rtfm / local docs skill if present known package/service: pinned local/project docs → official versioned docs → the known publisher's own read-only search/filesystem MCP when available → Context7 for cross-publisher library resolution, version selection and examples → Mintlify Index only as optional cross-product discovery after live quota proof → primary-page/source verification; unknown provider: API Finder only for candidate discovery → selected provider's official docs/terms/auth/rate limits/status → general WebSearch last A directory's auth, CORS, price, availability, or “open source” label is being treated as authority; the package/version/source revision is unknown; a publisher index excludes hidden/private material; Context7 has no matching version or collapses read-only docs search into an adjacent authenticated Admin MCP; Mintlify Index's shared 1000-request daily quota is unavailable for the run or it falls back to unbounded web results; or generated retrieval/example is treated as authority rather than a cited lead
Service/API access Agent-Native CLIs PrintingPress <svc>-pp-cli → print one → official CLI → MCP → raw HTTP IDE-only task needs MCP discovery; no CLI/auth exists
Project CLI and release engineering incumbent project scripts/workflow + service-cli-registry freeze release scope, source revision, targets, credentials, and rollback → keep a small project CLI on ordinary TypeScript; use pinned Stricli only when typed nested routing, isolated context, lazy commands, library mode, and shell completion are real needs → use the service's official generated client when it preserves typed errors/pagination and keeps credentials caller-owned → use pinned Craft only for a measured multi-target prepare/changelog/artifact-publication job → use pinned Fossilize only when Node SEA multi-platform binaries are the chosen distribution → use pinned binpatch only when measured binary-update bandwidth/latency justifies delta generation and safe application → preview prepare/changelog/target manifest → build/sign/SBOM/provenance → platform smoke tests → separately approve publication → verify release → preserve full-download fallback and rollback A five-tool stack is adopted wholesale; a small script gains a CLI framework; a generated client outranks an official CLI without a repeated API need; Craft is used for one simple target; SEA compatibility/signing/libc/assets are unproved; a delta updater can replace an executable without authenticated manifest, exact base/target digests, size/resource caps, cumulative verification, fallback, rollback, and explicit update authority; prepare or build success is treated as publish approval
MCP server or client implementation MCP as the Integration Standard + official versioned SDK freeze client/server jobs, transport, protocol era, capabilities, auth, tool/resource/task/app surface, authority, and compatibility matrix → prefer a current official SDK → 2026-07-28 stateless core for new HTTP deployments when every required client supports it → keep explicit application state in resource/task handles or owned storage → run official core, extension, auth, metadata, and back-compat conformance → preserve legacy 2025-11-25 negotiation only where needed “Stateless protocol” is mistaken for stateless application or safe tools; Apps/Tasks/auth support is inferred from a spec headline rather than client/SDK proof; legacy clients, issuer/scopes, credential isolation, cancellation, idempotency, resumability, observability, or mixed-version failure is untested
Portable skill plus MCP distribution source-owned skill/MCP directories first keep canonical procedures and server configs in their existing owners → use Agent Plugins v1 only when cross-client package portability is required → validate plugin.json, schema version, immediate-child skills/*/SKILL.md, optional version-matched mcp.json, plugin-root containment, environment placeholders, transport support, component diagnostics, license, signature/provenance, install diff, authority, doctor, and uninstall/rollback A working-draft package format is treated as a trust or marketplace standard; a plugin gains filesystem/network/tool authority from conformance alone; a symlink or relative path escapes the root; one failed component disables diagnostics or silently changes another; client-specific behavior is called portable
Application authentication / OAuth provider incumbent managed or framework-native auth route preserve an existing Clerk, Auth0, Supabase Auth, or project-native system → OpenAuth 0.4.3 when a team explicitly needs an MIT, self-hosted, standards-oriented centralized OAuth server across Node/Bun/Lambda/Workers and will own providers, user lookup/creation, token/password storage, email/PIN delivery, key rotation, abuse controls, recovery, audit, migration, and operations → prove authorization-code + PKCE, redirect/client allowlists, refresh/revocation, CSRF/state/nonce, multi-tenant subject claims, logout, recovery, backup/restore, and an independent security review before production “Self-hosted” is being treated as less security work; beta status, only six test files, or “mostly OAuth 2.0” is ignored; OpenAuth is assumed to manage users; password/PIN/email, KV consistency, signing keys, session revocation, account linking, MFA/passkeys, admin, audit, rate limits, or migration is unowned
Supabase optimistic/live client collections ordinary Supabase client + server/RLS truth ordinary queries/mutations first → Supabase TanStack DB adapter only for a React client that measurably benefits from typed collections, optimistic rollback, live queries, and Realtime reconciliation → keep Postgres/RLS/Auth authoritative, default Realtime off, test duplicate/out-of-order/missed events, optimistic conflict/rollback, reconnect/resync, SSR escape hatches, message volume, and spend cap The package's explicit experimental 0.0.1 state, absent SSR support, client-side fallback, Realtime message volume, RLS, conflict behavior, or production rollback is ignored
TypeScript AI application and harness surface Vercel AI SDK generateText/streamText/typed UI for app-owned loops → WorkflowAgent for durable resumable work → HarnessAgent for embedding Claude Code/Codex/Pi/OpenCode-class harnesses → freeze exact ai and subpackage versions, provider/gateway/model, approvals, context, files/skills, sandbox, telemetry, lifecycle events, auth, cost, deterministic tests, and UI states A release post is being treated as the API; provider portability is mistaken for behavior, policy, or data portability; an experimental harness adapter bypasses sandbox, identity, authority, resume, or eval proof
Typed effects / concurrency / resource lifecycle in TypeScript ordinary TypeScript plus narrow libraries ordinary async functions, result types, schema validation, and focused libraries first → Effect v3 stable when a codebase repeatedly benefits from typed failures, dependency layers, scoped resources, cancellation, structured concurrency, streams, schedules/retries, config, observability, and deterministic tests → evaluate Effect v4 beta only in a bounded branch with migration, bundle, type-check, runtime, learning-cost, interop, and rollback proof The team lacks a recurring cross-cutting problem; the paradigm would obscure simple code; main/v4 beta and v3/latest are conflated; an Effect skill or viral pattern is treated as proof of project fit
Sandbox provider selection incumbent governed execution plus Agent Sandboxes Match platform and workload → compare Vercel Sandbox, Cloudflare Sandbox SDK, Daytona, Runloop Devboxes, Docker Sandboxes, or Smol Machines → use ComputeSDK for an adapter layer → retain canonical rules, approvals, and proof outside the provider Documentation or a common API is mistaken for reproduced isolation, equivalent persistence, policy portability, or an adoption decision
Agent execution substrate trusted control plane + execution-tier receipt named capability, deterministic CLI, or isolated Code Mode for bounded work → agentOS virtual OS when shell/POSIX/process semantics are required → full sandbox/microVM/devbox/desktop for arbitrary native packages, Docker, browsers/GUI, heavy builds, GPU/hardware, kernel features, file watching, or higher-risk untrusted work → mount/escalate only the unsupported operation and preserve a content-addressed checkpoint The lighter tier changes program semantics or weakens the threat model; implementation/version, permissions, mounts, secrets, limits, state transfer, escalation reason, cost, or proof is missing; never run agent-generated build code as an ambient trusted-host process
Self-hosted E2B-compatible Firecracker fleet an already governed sandbox provider or dedicated Linux/KVM lab incumbent provider first → AgentENV only when E2B API compatibility plus self-hosted microVM fleet control is the measured need → pin and review install scripts/images → dedicated Ubuntu 24.04/Linux 6.8+ /dev/kvm host → loopback/private network or independent auth/mTLS proxy → tenant, image, egress, secret, quota, lifecycle, snapshot, encryption, restore, audit, upgrade, rollback, performance, and cost proof macOS or a general developer laptop is the target; upstream's missing authorization is ignored; the API is publicly exposed; privileged install code is piped blindly; tenant isolation, image provenance, snapshot custody, egress, deletion, recovery, or claimed sub-100ms timings are unproved
Managed persistent full-VM worker an already integrated full-worker provider existing E2B, Sprites, Dedalus, or Vercel route → Box by ASCII only when persistent Ubuntu, Docker-in-VM, SSH, desktop, dedicated IPv4, or template/fork lifecycle is the measured need → clean safeForThirdParties/noEnv environment from creation → external account bearer key → tenant/TTL/spend/start/concurrency/egress/teardown receipt → controlled API or pinned-SDK pilot A bounded or virtual-OS tier is sufficient; the opaque no-checksum CLI installer is piped into a shell; a previously trusted snapshot is assumed scrubbed; workers receive the account key or ambient credentials; public desktop/service access lacks separate approval; provider isolation/performance claims are unreproduced; the linked privacy policy remains 404 or written approval for third-party compute access is absent
Real-time collaborative coding room existing governed local/remote session and an intentionally shared worktree direct pairing in the incumbent project/harness → Jam only for a disposable fork or deliberately shared workspace after participant/repository manifest, least-privilege BYOK key, secret and egress boundary, spend limit, invite/revoke/delete test, and separate public-preview approval → self-hosted/team runtime only when persistent collaboration, RBAC, audit, backups, and operations are owned The room would inherit ambient private repositories or unrelated secrets; one shared terminal obscures who can observe/type/approve/spend/export/invite/publish; Jam server isolation, authorization, deletion, provider configuration, or recovery is unproved; a demo link is being treated as repository or publication authority
Shared watchable browser or disposable desktop ordinary screen share or governed browser session screen share for view-only help → Neko for an intentionally shared disposable container/VM when synchronized audio, WebRTC, clipboard, or multi-participant control is the actual need → isolate one task/profile, replace default credentials, terminate TLS, restrict member roles and clipboard/host authority, constrain egress and downloads, patch images, record joins/control transfers, test revoke/session expiry, then destroy the room Neko's default multi-user passwords (admin/neko) or broad host/clipboard rights remain; ICE/TURN credentials, ports, public IP exposure, browser profile, downloads, secrets, recording/telemetry, content rights, room deletion, or participant consent is unowned; “container” is being treated as sufficient isolation
Remote coding-agent control incumbent vendor remote surface, SSH, or private tailnet local control first → Tailscale/SSH for operator-owned private machines → T3 Code / T3 Connect only when its mobile/web/desktop control surface is the measured need and after exact relay build, OAuth/device identity, stored credential, environment link, tunnel, activity-publication, secret, repository, command, approval, disconnect, revocation, logs, update, and self-host fallback proof npx t3 connect is treated as a zero-trust boundary; a hosted relay receives more metadata or authority than intended; remote clients can execute or approve without a visible machine/repo/session identity; token loss, device removal, tunnel shutdown, activity publication, audit, or incident recovery is untested
Agent-operated real iPhone UI simulator/app-owned test API first app-owned XCTest/XCUITest or simulator route → official API/Shortcuts when the intended action has one → pinned phone-harness only for a user-authorized real-device flow where iPhone Mirroring is the necessary transport → pair interactively, grant the minimum Accessibility and Screen Recording permissions, verify --doctor, keep the mirroring window visible/frontmost, allowlist apps/actions, require approval for messages, purchases, account/security changes or irreversible effects, and verify each action from a fresh screenshot/OCR observation OCR coordinates are called a DOM; private notifications/photos/messages leak into traces; the device is locked or a different phone/window is targeted; focus or layout drift causes a wrong tap; setup-prompt text is granted instruction authority; no kill switch, permission removal, action receipt, or harmless canary exists
Authorized wireless-network audit approved network inventory and ordinary configuration review prove written ownership/scope and isolated lab target → non-disruptive passive discovery first → Airgorah only on Linux with dedicated monitor-mode hardware for the named authorized test window → deauthentication, handshake capture, or cracking requires a second explicit disruptive-action approval, containment, evidence handling, restore, and stop plan Ownership or written authorization is absent; neighboring networks/clients can enter scope; root, packet injection, deauthentication, captured traffic, credentials, wordlists, radio interference, retention, or legal requirements are unowned; the test can disrupt production or third parties
Multi-session agent terminal / operator cockpit incumbent Codex/Claude/T3/terminal surface incumbent surface for ordinary work → Termany only when five-plus concurrent sessions, nested panes, worktrees, remote hosts, inline diffs/files/browser, activity state, or cross-agent usage accounting is a measured advantage → desktop/local loopback first → review AGPL/commercial-license choice, bundled node-pty/Tauri server, transcript/session reads, SSH identity files, BYOK database, paste files, port forwarding, process termination, global hotkey, updates and eventual cloud/auth boundary Another cockpit duplicates existing operator state; local WebSocket/API is exposed beyond loopback; transcripts, keys, SSH credentials, files, remote ports, shell/process-kill authority, telemetry, or cloud roadmap are unowned; AGPL network use or commercial terms are ignored
Native terminal rendering / record-replay incumbent terminal emulator and PTY stack incumbent terminal for ordinary operator work → Native SDK terminal pattern only when a product needs an owned PTY plus libghostty-vt-style parser/render core, deterministic input/output recording, replay fixtures, and native UI integration → freeze terminal-core revision, byte stream, dimensions, locale, shell, escape-sequence corpus, redaction, and replay hash A demo is treated as a safe shell; PTY process authority, transcripts, secrets, clipboard, links, file transfer, input injection, terminal escapes, persistence, remote access, crash recovery, performance, accessibility, or platform support is unowned
Multi-agent organization / fleet coordination one named workflow in the current harness direct Codex/Claude/Hermes task for one bounded job → Orca-style interactive isolated-worktree supervision when parallel coding candidates and mobile/diff steering are the need → Paperclip only when 3+ persistent or heterogeneous agents need shared goal/task ancestry, schedules, budget attribution, approvals, workspaces, recovery, and an operator control room → keep Kevin-Wiki as canonical memory and require verified writeback The work is one task or one agent; an org chart adds ceremony; source master is not green; telemetry, auth/exposure, least-privilege skill/secret policy, workspace isolation, backups, upgrades, collision/retry recovery, costs, or canonical-memory boundaries are unproved
Human/team project and issue tracking incumbent repository issue tracker GitHub/GitLab issues for code-local bounded work → Plane when a team needs shared issues, cycles, modules, roadmap, docs, and multiple views with cloud or owned self-hosting → freeze Community-edition feature needs, AGPL obligations, identity/RBAC/integration/export policy, resource/upgrade/backup/restore owner, and one representative project migration before promotion One agent or one repo is the whole job; the tracker would duplicate canonical work state; Community and paid-edition features are conflated; team data, integrations, email, auth, backups, upgrades, export, or self-host operations are unowned
Multi-session decision discovery wayfinder ordinary bounded decision: grill-with-docs → clear current-context synthesis: to-prd → large/foggy route: manual /wayfinder with local map, decision tickets, acyclic blocking edges, frontier, fog, and one decision per session → when the route is clear, to-prd/to-issues → execution through workflow/goal contracts The work fits one session; tickets describe deliverables rather than decisions; the graph has cycles or stale frontier entries; fog is being invented into premature tasks; external tracker writes, access provisioning, spend, production, or subagents are implied without authority
Planning interview grill-with-docs map the whole decision-tree shape before depth → compute every unresolved decision whose prerequisites are settled → ask that frontier as one numbered round with recommendations → investigate discoverable facts directly → recompute after Kevin's answers → update CONTEXT/ADRs/prototype verdicts → exit only when the frontier is empty or deferrals have owners and Kevin confirms shared understanding Independent questions are serialized one per turn; dependent questions are asked in the same round; recommendations silently become decisions; a tree diagram alone is called alignment; downstream PRD/code begins before Kevin confirms the shared understanding
Session/phase transition agent-iteration-loop chooses route; handoff owns recipient/context transfer enough useful attention + bounded next step: continue → disposable context + independent next work: fresh session after durable writeback → valuable context crosses agent/harness/directory/project/human: handoff → bounded authorized AFK branch with parent synthesis: delegate → same owner and relevant context need a smaller working set: compact Session commands are assumed portable across harnesses; handoff becomes end-of-session ceremony; delegation broadens authority; clear/compact discards uncodified decisions; a context summary is mistaken for canonical memory
Domain identity and DNS delegation reputable registrar + named DNS owner production/auth/email/OAuth identity → paid registrar with recovery, renewal, transfer, DNSSEC, and named ownership → DigitalPlat FreeDomain only for disposable demos, learning, or non-identity-critical community projects with expiration monitoring and current-policy verification The name controls login, email, package identity, brand trust, recovery, or a production service; free namespace availability, renewal, or policy is being assumed from an old post
Disposable or catch-all test email inbox incumbent transactional-email sandbox or test mailbox provider sandbox for ordinary delivery tests → Maildrop for an owned lab domain that needs random/custom addresses, password-protected inboxes, automatic clearing, and a simple JSON API → containerize on a non-root internal port behind TLS/auth, define retention/deletion and abuse/rate limits, and keep sending disabled unless separately authorized The address is an identity/recovery inbox; the goal is mass account creation, spam, evasion, or impersonation; SMTP port 25/root exposure, catch-all privacy, shared password, public indexing, attachments, deletion proof, outbound mail authority, DNS reputation, monitoring, backups, or incident ownership is unresolved
Browser/computer use Computer Use and Browser Automation Patterns browser-harness → agent-browser CLI/MCP → Chrome DevTools MCP → cursor-ide-browser → screenshot/computer-use fallback Need deterministic CI, use Playwright first
Authorized protected-page retrieval Web Scraping Stealth (TLS Fingerprinting & Anti-Bot Evasion) prove authority and record terms/rate/data boundary → official API/export or direct fetch → managed extraction/crawl → ordinary browser control → classify transport/browser/protocol/IP/challenge failure → guarded Camofox Browser pilot only when real-Firefox shape and JS/property control match the failure → preserve target-specific quality/security receipt and fallback Authority is absent; account/login/access-control bypass is implied; the target works through a less evasive route; a promotional bypass claim replaces a controlled test; Camofox bind/auth/telemetry/profile/trace/network/package/browser gates are unproved
Personal data-broker removal Unbroker verify self-subject or signed scoped authority → California resident: official DROP directly, preserving DROP ID/status securely → read-only people-search gap map and official broker channels → assisted field-level plan → one canary submission → external receipt and re-scan → recurring read-only monitoring; use Unbroker's deterministic ledger only after version/license/storage/cloud/email/browser admission Authority is only a caller-supplied boolean; a bulk/blind opt-out is outside official DROP; encryption/key separation, PII processors/retention/recording, current jurisdiction/registry, per-broker recipe, composite license, live success/removal/relisting, or external filed/removed proof is missing; no government-ID, phone/fax/mail, solver, or stealth automation
Local dev server URLs Portless portless skill → portless doctor → framework-specific port/host flags → manual port routing last The project cannot use Node 24+ or local proxy/CA changes
Personal always-on server / appliance managed VPS or owned mini PC first managed service when operations should be external → low-power mini PC when stable updates, storage, networking, and battery-free unattended service matter → rooted Android/Termux phone only as a personal lab after exact device/OS, charge limit, thermal and battery-health monitoring, fire-safe placement, stable power, root/chroot trust, VPN/ingress, update owner, off-device backup/restore, restart and power-failure tests, and total-cost comparison Production, hostile multi-tenant, or irreplaceable data; PRoot/chroot is mistaken for isolation; shared Android kernel/network/UID, rooting, software rendering, power management, OS updates, swelling/fire risk, backup restoration, unattended recovery, or measured reliability is unresolved
Deterministic browser tests Complementary Browser Testing Playwright → browser-harness/agent loop to discover new assertions The question is visual judgment, not repeatable behavior
Before/after interface evidence Browser Testing Skills + Frontend Frontier Radar Workflow freeze before/after URL or fixture, state, viewport, selector and full-page policy → agent-browser diff screenshot for machine delta → pinned before-and-after only when a paired local image/Markdown/PR artifact helps human review → explicit approved upload adapter if publication is required → behavioral, accessibility and performance proof stay separate A screenshot pair is called correctness; authenticated/private captures use the package's default 0x0.st upload; before/after states differ in data, auth, viewport, timing or selector; PR mutation/publication lacks approval; PolyForm Shield terms are treated as permissive open source
Direct human artifact feedback human-review agent-run audit for independent diagnosis → human-review for a user-requested local HTML, rendered Markdown, or localhost edit/comment loop → map the exact batch to source → target-specific tests/build/render → acknowledge → separate approval for merge/publish/deploy A remote URL or untrusted page is opened; plain HTML is mutated without authority; rendered Markdown/HTTP DOM is written over source; exact user wording is paraphrased; a stale save wins; acknowledgement happens before full apply/proof; feedback is treated as approval
Open-weight / model choice The Eval Loop (Slop Is an Output Problem) + model tool page existing harness/provider → Vercel AI Gateway / AI SDK → OpenRouter Fusion for multi-perspective research/critique → captured model pages such as GLM-5.2 or Kimi K3 → Ollama/llama.cpp for ordinary local serving → KTransformers only for measured large-MoE CPU/GPU placement or fine-tuning needs No same-harness benchmark for the task; provider/privacy/license or local hardware fit is unproved; simpler serving works
Product and website analytics PostHog skill / PostHog Ecosystem PostHog for product events, flags, experiments, replay, LLM traces, warehouse, and debugging → Plausible for deliberately simple aggregate website analytics via managed service or operated AGPL community deployment “Privacy-first” or cookie-free is being treated as automatic legal compliance; product behavior/session replay is required; self-hosting cost, upgrades, backups, ingress, data policy, or EU/provider requirements are unowned
SEO/GEO measurement ai-seo + seo-audit actual audience/channel usage and business goal → Google Search Console dedicated generative-AI report when available plus overall Web report → first-party referral/conversion analytics → repeated cross-product prompt panel with query/model/date/locale/account/mode and cited URL → crawl/index/Core Web Vitals proof → corroborated backlink/organic evidence → third-party tracker only as a method-disclosed observation; TraceDR only as an optional 24-hour Ahrefs Domain Rating trend signal A mention, citation, impression, referral, and conversion are collapsed into one “rank”; a single answer screenshot or vendor metric is treated as stable/causal; a third party claims access to internal Google ranking/AI systems; training crawlers are confused with search crawlers; first-party data or reproducible panel metadata is absent
Workflow and integration automation code + executable skills/workflows reviewable script/skill/workflow for canonical Kevin-Wiki loops → project-scoped n8n when visual operations and its connector catalog are materially useful → Gumloop when a managed AI-first canvas and hosted connectors/models are desired A visual canvas would become the only canonical definition; credentials/webhooks/egress, idempotency, retry, export, backup, or upgrade behavior is unproved; n8n fair-code/enterprise license boundaries are being called blanket open source
Shopify theme and store work shopify-commerce + Shopify freeze repo/store/theme/data/authority → read brand, product, and current theme facts → official docs/schema validation → Shopify CLI development or unpublished theme → section-by-section Liquid/JSON implementation → Theme Check + exact-store preview + mobile/a11y/commerce/performance proof → explicit approval for Admin mutation or live publish; use Shopify Hydrogen only for headless React storefronts and Storefront MCP only for customer commerce A generic MCP receives store credentials; Dev MCP is mistaken for Admin access; the live theme is edited during exploration; third-party skills lack license/audit; customer data enters prompts/telemetry/receipts; merchant editability, cart/variant states, rollback, or publication approval is missing; a screenshot or subjective AI score is called CRO proof
Conversational forms / surveys incumbent product form or managed form service keep project-native forms for bounded flows → managed service when operations should stay external → Coder Apps Bundle HeyForm when self-hosting or response-data control justifies owning uploads, mail, abuse prevention, backups, upgrades, and AGPL-3.0 obligations
Self-hosted availability / infrastructure monitoring incumbent provider plus application-specific observability provider health checks and existing uptime first → Sentry for application exceptions/traces → Coder Apps Bundle Checkmate for owned HTTP/ping/TCP/gRPC/DNS/SSL, hardware, incident, and status-page monitoring
Application observability coverage audit runtime telemetry plus framework-native tests verify real traces/logs/audits first → use evlog map for a supported Nuxt/Nitro/Next.js/TanStack static inventory and explainable gap score → inspect sensitivity reasons and disabled checks → emit JSON; use --baseline as a reviewed lockfile-like ratchet or --min-score as a floor → correlate the highest-risk routes with runtime canaries and incident queries A static 100 is called correct telemetry; unsupported framework or deep logger indirection is ignored; a heuristic money/auth/PII classification grants security authority; disabled checks hide missing events; generated suggestions are applied without route semantics, privacy, volume, or runtime proof
Application deployment / self-hosted PaaS Coolify working managed incumbent when low operations is the goal → Coolify when owned servers, Docker portability, integrated apps/databases/services, previews, and predictable infrastructure spend are explicit requirements → team-scoped MCP with expiring read token for observation → frozen UUID/health/deployment/backup plan → separate short-lived deploy token for approved lifecycle actions or official CLI/API plus explicit write authority for configuration → canary → health/log/DNS/TLS/backup/restore/rollback proof → canonical writeback The decision is based only on stars or a provider bill; no one owns root-level SSH reach, control-plane isolation, OS/Coolify updates, ingress/API exposure, secrets, backups and restore, monitoring, capacity, incidents, or total operating cost; routine agents receive read:sensitive, write, or root; MCP's “read-only” label is trusted despite its documented lifecycle tools
Application deployment / managed PaaS incumbent project platform keep the existing Vercel, Cloudflare, AWS, or project route when it works → Coder Apps Bundle Sevalla when one managed app/database/storage platform and lower infrastructure labor are measured requirements → official CLI/API before hosted MCP for repeatable mutations
Standalone Markdown documentation site incumbent docs framework + agent-docs preserve an existing Docusaurus, VitePress, MkDocs, or app-native site → Coder Apps Bundle DocMD when a new site specifically needs static output, local search, llms files, stdio MCP, Mermaid, or an OKF graph projection with minimal framework code
Agent-maintained code or personal wiki projection project-owned AGENTS/SKILL/docs plus agent-docs repair sparse project instructions first → OpenWiki only when one repository or bounded personal-source set needs agent-generated linked Markdown, connector synthesis, managed pointer blocks, an update PR loop, graph visualization, or OKF export → pin version/revision, preserve user instructions, scope sources/credentials, opt out or record telemetry, sandbox provider code, diff/no-op updates, cite source files/traces, test private ignores/deletion/drift, and keep rollback Kevin-Wiki's canonical brain or an incumbent docs system is being replaced; generated prose is treated as source truth or human intent; global install, connector OAuth, provider keys/tracing, default telemetry, CDN-loaded visualizer, scheduled writes, secret files, generated diagrams, or update PR authority is unreviewed
Collaborative product design existing Figma files + Figma skills/MCP Figma for incumbent files, libraries, plugins, Code Connect, and MCP workflows → Penpot when open SVG/CSS/HTML/JSON standards or self-hosting/governance materially outweigh migration cost The task is only code-side interface implementation; format fidelity, library migration, plugins, fonts, auth, backups, collaboration performance, or self-host operations are unproved
Team workspace / editable knowledge wiki/qmd for Kevin's canonical brain wiki/qmd for durable agent knowledge → existing company/project workspace → AppFlowy only for a bounded collaborative workspace whose export, backup, auth, upgrade, and AGPL obligations are owned A second workspace would become ungoverned canonical memory; migration/export fidelity, mobile/client behavior, sync, recovery, or operational ownership is unproved
Booking and scheduling surface Google Workspace Calendar / managed booking existing calendar skill and managed booking first → calcom/cal.diy only for a personal, non-production self-host experiment with explicit server, database, security, mail, backup, and upgrade ownership A production/team booking surface is needed; upstream's personal/non-production warning is ignored; calendar credentials, notifications, availability correctness, abuse controls, or uptime is unproved
Password and secret custody existing organization-approved password manager / platform secret store current approved manager and OS/cloud secret store → official Bitwarden service/clients or exact-module-reviewed self-host only when migration and operational ownership are deliberate Source availability is treated as package safety; the exact Bitwarden module/license is unknown; recovery, MFA, emergency access, backups, updates, server exposure, client trust, or npm supply-chain history is unreviewed
Agent credential mediation existing secret store plus narrow first-party OAuth/service credentials raw secret stays outside model/sandbox → scoped proxy or short-lived capability with provider, tenant, audience, tool, expiry, revocation, rate/cost and audit enforcement → LangSmith Auth Proxy, DAuth, or Treg only after exact service/terms/threat-model proof "Agents never see keys" is accepted without proving confused-deputy, response leakage, logs, tenancy, server identity, scopes, revocation, outage/recovery, or operator access
Agent skill supply-chain scan skill-auditor plus source/manual review pinned tree and dependency/global-write diff → NVIDIA SkillSpector static scan and optional semantic checks → repository tests/CI → manual trigger/behavior/security review → isolated pilot → held-out value proof A clean scanner result becomes installation approval; LLM semantic scanning leaks proprietary skill text; baselines suppress new findings; behavior/dependency/global-write review is absent
Research experiment execution Research Experiment Workflow freeze baseline/evaluator/held-out set → primary literature and attributed recipes → data audit → exact-source sandbox preflight → one pilot → monitored run/ablation → held-out confirmation and writeback; use ML Intern - Hugging Face Research Agent conditionally for HF-native literature/data/Jobs work, pinned Dive into LLMs only for 11-theme curriculum/recipe discovery followed by current-source reconstruction, AutoResearch (Karpathy) - Applied to Marketing for a bounded fixed-scorer code ratchet, and CoTCodec for orchestration-variable studies The task is only information gathering; no discriminating evaluator exists; data/license rights, compute budget, trace egress, or mutation authority are unresolved; do not infer notebook runnability from stars/free access, and do not use ML Intern headless/local-default execution on sensitive work
Self-hosted research workspace Self-Hosted Research Workspaces wiki/qmd for canonical Kevin memory → Open Notebook for a bounded source-grounded project notebook and exportable citations → Sevenfold for Kevin's academic-research product context → Odysseus only for a hardened, isolated full operator-workspace evaluation The task is only retrieval; a second application database would become ungoverned memory; credentials, auth, provider egress, backups, citations, export, or broad shell/email/MCP authority are unproved
Code review / bugs gstack-review gstack-review → counterfactual → cross-modal-review → language/security-specific skill The user asks for QA, design review, or CI repair instead
QA / dogfood gstack-qa gstack-qa → dogfood → beta-dx-walk → browser-harness / agent-browser Need deterministic regression, write Playwright
Unknown UI component or style name ui-vocabulary local UI Component Vocabulary clue card → generic public NameThatUI search only when the literal query contains no private project/customer/repository data → verify the candidate against Apple HIG/Developer Documentation, WHATWG, WAI-ARIA/APG/WCAG, MDN, or the owning framework → Component Gallery for cross-system implementation comparison → Component Library Sources for source selection The component is already named; the task is only implementation or polish; a screenshot is being used to infer unproved focus, modality, keyboard, dismissal, or selection behavior; an external query would disclose private language; the discovery result lacks primary-source verification
Design/frontend taste-skills-playbook taste router → classify blank slate/redesign/addition/refinement/authorized reconstruction → local system constraints → Impeccable v4 direction board + intent contract for blank slate/redesign; evaluate Hallmark instead when theme/macrostructure exploration, design-DNA study, or its audit vocabulary is the distinct need; use incumbent-preserving local polish for additions/refinements; use Website Cloner Skill only for an authorized live-reference reconstruction with DOM/CSS/behavior extraction and independent accessibility/rights gates → browser verification The task is code correctness rather than UI quality; an external catalog is being treated as a house design system; reference-site rights or identity are unclear; do not stack competing rule systems globally
Expressive animated component source animated-component-libraries + Component Library Sources name the exact effect → reuse an owned primitive when available → search Originkit or the ranked source shelf for one item → freeze preview, catalog metadata, MCP/CLI client, exact delivered payload, component terms, dependencies, remote imports, and composed items separately → fetch into a clean branch/worktree without dependency scripts → retokenize → prove semantics, keyboard/touch, reduced motion, teardown, visibility/offscreen suspension, static fallback, mobile behavior, and measured frame/GPU/battery cost The request is vague “premium” styling; bulk installation or authentication is proposed; catalog count/engagement/copy affordance/marketing label/client license is being treated as source authority; exact component rights or payload are unavailable; several spectacle effects would replace the product's design language
Creative/design application builder Toolcraft when a real canvas-and-controls product is the job name the product output and reference/rights boundary → confirm that canvas, meaningful controls, and upload/history/pan-zoom/layers/timeline/export needs justify an editor architecture → pin Toolcraft source and generate project only after explicit selection → preserve signed runtime versus product-source boundary → declare interaction, animation, and artifact intent → map visible entities and reference behavior to functional/browser/export proof → use Fieldwork — Visual Design IDE for later UI inspection and HyperFrames when the deliverable is deterministic video The job is an ordinary website, dashboard, form, one component/effect, or small settings overlay; interaction/export intent is unresolved; generated dependencies/source custody or protected verification cannot be owned; a demo reel or popularity signal is the only evidence
Interface icon system frontend-design-taste + UI Design Defaults (Cheatsheets) inspect incumbent imports and product jobs → name a concrete failure → render 20–30 real concepts at 16/20/24px → compare one family/style challenger → verify exact revision and license → owned AppIcon wrapper and semantic map → direct imports → accessibility/RTL/state proof → production-bundle comparison → keep or roll back; route brand/company marks separately through The SVG “Lucide is common” is the only reason; packs are mixed per component; the exact free/Pro/set license is absent; Solar rights remain unresolved; wildcard imports or package-wide runtime lookup lack output proof; a visual alternative does not improve the product
Raster/photo editing existing licensed editor or GIMP preserve the source and required color/format metadata → existing project editor first → GIMP 3 for open-source desktop editing → PhotoGIMP only when a human specifically benefits from Photoshop-like shortcuts/layout, after backing up the complete GIMP configuration The task is generation rather than editing; a config patch would overwrite a tuned GIMP profile; automated use lacks a deterministic export and visual-proof loop
Graphics/diagrams create-graphics define communication job → exact HTML/SVG for trusted structure → Excalidraw for editable whiteboards → Mermaid only when sufficient → image generation for visual ideation → target-size/accessibility proof The task is interactive analytics; use Plotly/Evil Charts instead
Pre-existing illustration / stock asset create-graphics + Design Inspiration Galleries use the shelf only for discovery → exact asset, creator and primary distribution → current terms/plan and local hash → rights class, attribution, count/price, modification, client-transfer, redistribution/trademark/ML checks → target-product proof and retained rights receipt “Free,” zero price, a roundup, preview or download button is treated as a license; rights are unresolved; the asset becomes the logo/identity, standalone product, catalog, model-training input or redistributed source without explicit authority
Chinese/Japanese calligraphy artifact hosted 墨 Ink / Ink Inky as a bounded authoring tool choose Chinese or Japanese data → write or trace → set guide, stroke order, meaning, orientation, surface, weight, wetness, speed, formality and seed → preview write-on/trace behavior → export PNG/JPEG/SVG/MP4/WebM/GIF or a self-contained React canvas component → inspect the literal output and retain the settings/source post Open-source or redistributable source is required; the artifact needs editable vector paths beyond the tool's SVG output; exact font/calligraphic authority is being inferred; hosted-runtime visibility is treated as a license grant; accessibility, cultural/linguistic review, browser encoding, output rights, or target-product fit is unresolved
Recurring-character editorial illustration illo explicit /illo or identity-consistent editorial-art request → doctor → source/artifact-job/thesis lock → frozen character reference on every render → QA-passed set anchor → re-render failures → manifest and final-delivery proof The request is generic image generation, photography, a logo, product UI, an exact formal diagram, or deterministic iconography; never spend through a fallback provider without explicit approval
Web-to-native shell PWA or Native SDK (formerly Zero Native) PWA/installable web first → Zero Native/Tauri/native shell for a project-owned cross-platform app with narrow bridges → WebToApp only for an isolated Android on-device APK workshop and owned test content Broad Android permissions, runtime downloads, cleartext/MITM/CORS bypass, userscripts/extensions, repackaging, signing-key custody, or store-policy/data-safety review is unresolved
Self-hosted app deployment platform existing project deployment provider and IaC prove the repeated deployment pain and target topology → compare Coolify/current incumbent with OpenShip on auth/RBAC, SSH and Git credentials, build isolation, artifact provenance, secrets, networking/TLS, databases/backups, logs/metrics, updates, rollback and total operations → isolated non-production canary A dashboard or "own infrastructure" claim substitutes for control-plane security and operational ownership; MCP/agent access lacks verb-specific approvals and receipts
Charts/data visualization Plotly + Evil Charts + Flint with flint-chart-author when agents author chart intent analytical question or notebook → Plotly for Python/JS interactive analysis or Dash; product dashboard with owned React treatment → Evil Charts/Recharts; agent-authored, human-editable semantic chart request that must compile across Vega-Lite, ECharts, Chart.js, Plotly or Excel → pinned MIT Flint ChartAssemblyInput through the installed source skill, with the host retaining data transforms, row binding, field/policy validation, storage, UI controls, local-file policy and backend selection → browser/accessibility/export verification A native table or static editorial graphic is clearer; an agent-generated backend blob is stored instead of the small semantic input; Flint is treated as ETL, renderer, policy owner or chart-state database; Python packaging, target backend, field identity, malformed data, accessibility, export fidelity or target-product proof is missing
Technology vitality Is This Tech Dead? as first-pass signal aggregate score → official release/security/maintainer evidence → project constraints → migration cost A dashboard score is being treated as a final adoption verdict
Agent UI / messaging surfaces Agent GUIs CopilotKit / AG-UI Protocol for app-embedded run state and approvals → Page Agent only when a product-owned page needs a client-side natural-language GUI agent over its text DOM → JSON Render -- Generative UI Framework or OpenUI -- Generative UI Language and Runtime for catalog-constrained generated UI → Photon Spectrum for a messaging-provider adapter → AstrBot only when the product needs a full self-hosted multi-channel bot, model/MCP/skill/knowledge/plugin runtime, and sandbox → Human-in-the-Loop Control before side effects The product is not exposing an agent interaction surface; external browsing or CI is the job; API keys, demo-CDN terms, prompt injection, mutations, AGPL, or marketplace plugins are unreviewed
Data movement Agent-Native CLIs ingestr → service pp-cli / official CLI → custom dlt/SQLAlchemy/pandas Connector missing or transform is domain-specific
Document/PDF work Retrieval Quality + artifact skill uncertain or mixed corpus: pinned local pdf-inspector preflight → retain class, confidence, page-level OCR route, and reason → read/extract: liteparse first for the supported local corpus; compare pinned AnyDoc when mixed legacy Office (.doc/.ppt/.xls) coverage, one document model, Rust/Node/Python/WASM bindings, or typed malformed/resource-limit errors are the measured gap → markitdown → mineru → omniparse by complexity; DOCX/XLSX/PPTX create/edit/render: documents, spreadsheets, or slides skill → OfficeCLI when a single-binary command surface and HTML/PNG render loop are useful; PDF mutate: PDF Operations → PDFCraft for browser-local one-off/visual workflows or Stirling-PDF for self-hosted API/OCR/batch service → preserve original plus rendered/text/page-count/hash comparison A project benchmark with a private 100-document corpus, AnyDoc-as-reference trigram score, one warm conversion, or different supported-format sets is called independent proof; scanned/image-only PDF, page geometry, screenshots, OCR, or layout coordinates are required; a classifier confidence score replaces representative visual/text checks; page numbering or OCR reasons are dropped between bindings; the task is a web page, use defuddle; malformed/encrypted/large/form/annotation/signature/redaction fidelity is untested; AGPL/network-deployment obligations are unowned; do not pipe remote skill files into an agent or let officecli install modify every detected harness without source and diff review
Authored and assembled video generation HyperFrames Active /hyperframes router + pinned strict deterministic fixture for bespoke agent-authored launch/demo/explainer/PR/social scenes → approved local media routes for user-supplied assets while upstream media-use remains in progress → MoneyPrinterTurbo for approved-script, high-throughput stock-footage voiceover shorts under Content Pipeline Workflow → Remotion only for an existing React composition system after its custom-license check → Remocn only as a component catalog inside that approved Remotion route → Text-to-Lottie for editable vector animation → AI video MCPs only for stochastic footage The user needs live UI/browser verification; source/claim, media/music/voice rights, provider credentials/telemetry/cost/egress, complete artifact QA, license eligibility, or explicit publish authority is unresolved
Stochastic media generation task-specific approved provider existing narrow provider/API skill such as Sume or Higgsfield MCP → OpenRouter/image-generation route → Open Generative AI only as a human UI/local-inference reference or after explicit MuAPI egress/cost review; audit one recipe rather than bulk-installing Generative Media Skills Deterministic scenes are required; subject consent, source rights, identity manipulation, cost, provider egress, or output provenance is unresolved
Article and spoken-content illustration approved image-generation route + Content Pipeline Workflow extract narrative beats and visual promises → write a shot/config strategy → use the existing image-generation/illustration skill with explicit style, ratio, count, title, reference, rights, and QA → use GiMi's illustration skill as a retained pattern source for strategy-first article illustration, custom-character enrollment, calibration, and paired output receipts → use only IP=none or Kevin-owned character/style references unless separately licensed The Gimi character assets are assumed MIT; a saved style becomes universal; text-only prompting is claimed to preserve character identity; provider/output rights, reference consent, disclosure, cost, placement, accessibility, or complete visual QA is missing
Marketing/content content-strategy / social-draft product and audience foundation → when positioning or market validity is uncertain, run the Content Pipeline Workflow market-research lane over a dated mixed corpus and extract assumptions, falsifiers, contradictions, adversarial case, and customer-interview questions → content-strategy or the needed local SEO/copy route → for multi-specialist campaigns, reuse the Eve pattern of one non-authoring lead, fresh bounded specialists, one singly-owned brand context, least-capability connectors, and exact irreversible-action approvals → use Origami only as an optional current-terms-reviewed list/enrichment provider with no-send default → social-content for channel adaptation → social-draft + Kevin voice gate → performance evidence; use coreyhaines31/marketingskills and charlie947/social-media-skills as reviewed source collections for a distinct missing job, never as automatic bulk installs Synthesis is mistaken for market truth; the corpus is stale, homogeneous, or entity-misaligned; research silently becomes scraping, enrichment, outreach, spend, account mutation, or publishing; Origami source accuracy, credits, terms, legal basis, suppression/opt-out, export/deletion, or credential custody is unresolved; obtain exact approval and use the service-specific write boundary
Programmatic SEO page families programmatic-seo + Writing and Content Skills demand and domain-risk review → versioned niche/data taxonomy → strict payload schema → validator → page-type renderer → similarity/factual/raw-response/rendered proof → staged noindex cohort → explicit index allowlist → Search Console and business-outcome monitoring → refresh/consolidate/remove; route an existing crawl/index/manual-action problem to seo-audit Page count or early traffic is the success metric; generation automatically changes index state; schema-valid output is treated as useful/correct; raw deep URLs have generic metadata/canonical/content; source rights, maintenance owner, index budget, cohort evidence, manual-action check, or removal path is absent
Career career-ops career-ops → recruiter-discovery/content-pipeline workflows The task is general content or marketing rather than Kevin's career system
Financial-market research terminal authoritative primary filings/data plus read-only analysis official filings and owned account statements first → Gloomberb 0.10.4 for keyboard-driven research, quotes, charts, filings, macro data, watchlists, notes, and structured JSON/CSV/NDJSON exports → pin providers and freshness, record source timestamps, use --dry-run where supported, and treat plugins/AI/cloud/broker connections as separate trust boundaries A quote, rating, AI screen, plugin, or cached provider output is treated as investment advice or authoritative/current; broker credentials, trade authority, --yes, account/profile binding, data terms, attribution, mutation receipt, or professional review is unresolved
Finance / small-business / legal operations professional owner plus a bounded workflow Claude Cowork's Anthropic-verified Finance, Small Business, or Legal plugin may be used only in that supported surface; otherwise route through the relevant spreadsheet/document/service capability and a locally documented workflow. Require source data, reversible drafts, human approval before money/customer mutations, and qualified finance or licensed legal review before reporting, filings, contracts, or compliance decisions A plugin is unavailable in the current harness, connectors are unaudited, source records are incomplete, or the output would be treated as professional sign-off
High-stakes model decision control Human-in-the-Loop Control typed authority/risk/evidence validation → deterministic deny/defer/escalate gates before model access → minimize/tokenize sensitive fields and fail closed on detector/residual uncertainty → freeze model-visible input and candidate set → model assists only inside the allowed decision vocabulary → human approval where unique authority/accountability remains → decision/resume receipt Policy prose is the only enforcement; raw sensitive data reaches the model before gating; invalid parse or ambiguity silently becomes approval; banking/research thresholds are copied without domain validation; no proof says whether the model ran or what the human approved
Memory/codification No One-Off Work update wiki → create/extend skill → automation if recurring → doctor/index/qmd; for session-derived skill improvement use SkillClaw only as a capture/proposal candidate under Controlled Skill Evolution - GEPA and SkillOpt and Skill Sync Workflow evaluation/promotion gates The task is genuinely one-off session context; captured sessions lack consent/redaction; a candidate would overwrite canonical skills without held-out proof and rollback
Agent loop design Loopy loop-me → loopy → skill-creator/automation; add agent-eval-library when the judge matters The task is a one-time plan, use borrowed-intelligence/improve
Agent self-improvement evals Agent Self-Improvement Eval Library receipt suite → npm run agent-self-eval → Red Queen Gödel Machine hardening → Loopy loop The task is ordinary content/product output quality; use The Eval Loop (Slop Is an Output Problem)
Guardrail-policy optimization agent-eval-library + Agent Self-Improvement Eval Library freeze target/harness/judge/splits/budget and one mutable policy → record unguarded, incumbent, and candidate attack plus benign/task metrics → accept only strict safety improvement inside the utility floor → preserve rejection and restore incumbent → start a new lineage if the evaluator boundary changes All-refusal wins; the unguarded baseline is absent; judge/suite/target changes mid-run; a deterministic single-turn stub is presented as production safety; tool/file/multi-step limits are omitted

Browser and Computer Use

The browser stack is split by task type, with one default: agent-browser owns agent-driven browser work. Playwright owns committed test artifacts.

Task Use first Why
Agent-driven browsing / scraping / form fill / screenshot / UI QA Browser Testing Skills Compact snapshots, refs, screenshots, sessions, batch commands, Chrome profile support
Logged-in Chrome or personal-account flow Browser Testing Skills with Chrome profile / auto-connect Same agent-browser workflow with real auth state
Simultaneous long-lived authenticated work with explicit agent/user ownership and takeover ego-lite only as a guarded dedicated-profile route after Browser Testing Skills Task-space ownership and handoff are the distinct capability. Audit the exact open harness and closed browser binary; use a dedicated profile/account, import minimum state, keep CDP local, and record ownership, mutations, and closeout.
Repeated authenticated website/Electron adapter Browser Testing Skills first; OpenCLI only as a guarded project-scoped candidate Prove the repeated route first. If adapting it, isolate a Chrome profile/project; inspect exact adapter, manifest, extension permissions, cookies, dependencies, doctor, and tests; require explicit authority for mutations.
React/perf/Web Vitals Browser Testing Skills Web Vitals and React render introspection
Visual bug hunt / responsiveness Browser Testing Skills diff screenshot, diff snapshot, viewport sweeps, rendered proof
CI E2E / repeatable regression Browser Testing Skills Deterministic, reviewable, gates every PR
Screenshot-first/self-healing real-Chrome edge case Browser Harness - Self-Healing Browser Automation Specialist fallback when compositor-level clicks or self-healing helpers are the point
Computed CSS / box model deep dive Chrome DevTools MCP Explain why the bug happens after agent-browser finds it
IDE-native quick look cursor-ide-browser Built into Cursor's vision pipeline
Desktop/Electron app electron skill, then screenshot/computer-use fallback Structured app automation before raw screen driving
Repeated website workflow Autobrowse (Browserbase Skills) / Browserbase skill family route Learn once, save playbook, amortize browsing cost
Undocumented website API discovery Browser to API candidate route; HAR-derived client only after a repeated need Start with official API/export and terms → capture only traffic Kevin is authorized to observe → isolate profile and account → redact cookies, tokens, PII, bodies, and signed URLs before durable storage → derive a read-only schema/client hypothesis → pin host, route, method, headers, auth/session lifecycle, pagination, rate limits, and response samples → add drift fixtures and fail closed → mutations require a separate target-bound authority receipt. A generated client is a reverse-engineered hypothesis, not provider support.
Cloud browser session with recording/live takeover/WebMCP Cloudflare Browser Run only for a declared remote-browser requirement Freeze account/region/pricing, connected profiles, session lifetime, recording opt-in/retention/delete/share, human-takeover authority, secrets, egress, and mutation receipts; ordinary local QA stays on agent-browser
MHTML capture for a DOM-dependent page Browser-owned controlled capture after ordinary HTML/snapshot fails Treat MHTML as a sensitive archive: it may embed DOM, inline resources, URLs, tokens, user data, and rendered state. Capture on an isolated profile, hash/redact/encrypt, restrict retention, and verify resource/fidelity gaps before parsing.
Share a local dev server privately Portless with Tailscale Serve / tailnet route Bind the intended app only; verify tailnet ACL/audience, generated certificate, route cleanup, secrets/test data, logs, and stop/revocation. --funnel is public exposure and needs separate approval.
Public localhost tunnel Tailscale Funnel or tunnelto only for an explicitly public, disposable target Add application auth, narrow routes, no real secrets/private data, expiry, monitoring, rate/abuse controls, teardown and external reachability proof. A random URL is not access control.
Download public web video for an authorized task platform export/API first, then yoinks/yt-dlp only for rights-cleared material Record source URL, owner/license/permission, platform terms, requested format, binary/dependency revision, output hash, attribution and deletion/retention. Never bypass paywalls/DRM/access controls or treat downloadability as reuse rights.
Fast read-only public-repository conversation without cloning Gitinspect as a convenience surface; local rg/qmd/Graphify remain the durable routes Gitinspect is a browser-local pi + just-bash virtual-filesystem pattern, not a source of truth or a replacement for pinned repository evidence. Do not place private repositories, provider keys, or GitHub tokens into an unreviewed hosted surface.
Exact-element feedback on a generated HTML artifact Lavish as a local, bounded review candidate Keep the core loop local. Export/share is a separate external effect: inspect remote assets, redact file paths and secrets, and require explicit approval before publishing to ht-ml.app.
Harness-specific browser side panel for Hermes Hermes Browser Extension only when Hermes session/context handoff is the job This is an adapter to the Hermes runtime, not the cross-harness default. Review extension permissions, local endpoint exposure, session data, screenshots, and model/provider changes before installation.
Small local HTTP browser-control daemon PinchTab as a specialist fallback after Browser Testing Skills Keep server, dashboard, MCP, and remote CLI on loopback/private networks. Profiles and tabs contain privileged state; public or multi-tenant exposure requires authentication, TLS, endpoint reduction, isolation, and an explicit threat model.

Search and Research

Need Use Notes
Search Kevin's compiled brain QMD - Local Wiki Search Engine Brain-first. Exact names via qmd search; synthesis via qmd query.
Ask codebase topology/path/blast-radius questions Graphify / graphify-sidecar Sidecar-first for focused graph context, then verify with source reads.
Research what people currently say last30days Multi-platform, engagement-weighted current signal.
Read/search live social/web/video platforms Agent Operations Skills Use when the source is X/Reddit/YouTube/etc.
Work interactively over a bounded private corpus with source chat, transformations, notes, podcasts, and an API Self-Hosted Research Workspaces Open Notebook first; export durable sources, citations, and conclusions back through source compile. Odysseus only for a separately approved high-authority operator-workspace pilot.
Build or enrich structured public-web lists Exa Agent Async research runs with outputSchema, grounding, cost accounting, and continuation; overkill for one search.
Current library/API docs NIA Docs - Docs as Filesystem for Agents Mount docs as files, then tree / grep / cat.
Find agent skills find-skills Search installable skills; run skill-auditor before installing.
Exact repo text rg Local repo search still beats every semantic layer for exact strings.

Model and Agent Runtime Choice

Do not pick a model because it is new. Pick the runtime path from the product constraint.

Constraint Route
Coding inside an existing harness Use the harness-native model picker first: Cursor / Claude Code / Codex / agent runtime.
AI feature in a Vercel/React app Vercel AI SDK + AI Gateway; evaluate model swaps behind the same interface.
Need one local endpoint over many operator-owned providers, subscriptions, quotas, fallback policies, or compression experiments Evaluate OmniRoute only after comparing AI Gateway/OpenRouter/LiteLLM and reviewing credential encryption, provider terms, outbound routing, remote tokens, MCP authority, compression fidelity, TLS impersonation, and optional MITM behavior. Keep loopback-only and least-authority surfaces until proven.
Deep research / architecture critique needs multiple model perspectives OpenRouter Fusion as alias/plugin/server tool; use selectively, not as a routine coding model.
Long-horizon open-weight coding/model eval Compare GLM-5.2 and Kimi K3 against the incumbent under one frozen task set, harness, tools, context policy, effort, budgets, and verifier. Record score, failures, retries, tokens, wall time, cost, and variance; promote only the measured model/provider/effort/harness tuple. Use a gateway or official hosted API for the bounded pilot; self-host only when data control or architecture research justifies hardware, engine, operations, privacy, and license review.
Native multimodal, million-token, long-horizon agent eval Kimi K3 is the captured candidate. Preserve its complete assistant messages, reasoning_content, tool calls, effort, and context policy; do not promote it from mixed-harness vendor charts.
Need a very large MoE model whose weights or experts do not fit ordinary GPU VRAM Evaluate KTransformers for CPU/GPU heterogeneous inference or LLaMA-Factory SFT only against the exact model, quantization, CPU instruction set, RAM, GPU, concurrency, and latency target; reproduce upstream benchmarks before adoption.
Need a self-hosted agent chatbot across many IM platforms with its own knowledge, skills, MCP, plugins, web UI, and code sandbox Evaluate AstrBot; audit AGPL obligations, platform credentials, sandbox isolation, model/data egress, and every community plugin. Use Photon Spectrum when only the provider/channel adapter is needed.
Need durable agent sessions / sandboxed filesystem runtime Vercel Eve, AI SDK HarnessAgent, Agent Machines, or Dedalus Machines - How They Work depending on product boundary.
Need a persistent personal operator stack on a local Mac Hermes Mac Mini Agent OS: Hermes + Hindsight + Honcho + qmd/Obsidian + agent-browser + ACP workers + Telegram threads. Treat Hermes runtime upgrades as environment migrations with config/profile diff, doctor, permission/privacy and regression receipts. For external lifecycle integration, prefer signed webhooks or A2A behind signature/event-dedupe/least-authority gates; keep Kevin-Wiki as canonical writeback.
Need coding-agent product comparison Compare ZCode, Vercel Eve, Loop, Cursor/Claude Code harness pages.
Need safety or content quality Build an The Eval Loop (Slop Is an Output Problem) first; do not substitute model vibes for output tests.
Need an agent to benchmark itself Use Agent Self-Improvement Eval Library first; keep the evaluator evolving with Red Queen Gödel Machine.

Service Access

For external apps and services, the hierarchy is interface-first:

  1. Existing PrintingPress pp-cli - local mirror, compound commands, low token cost.
  2. Print one - if Kevin will hit the target again and no pp-cli exists.
  3. Official CLI - gh, stripe, vercel, aws, supabase, etc.
  4. MCP - best for IDE/desktop agents, tool discovery, or services with no CLI.
  5. Raw HTTP - last resort; codify into a CLI/skill if repeated.

Headless auth: use Agent Cookie - Session State Sync for the Agent's Second Mac only after explicitly surfacing the security tradeoff.

For websites with no published API but observable XHR/fetch traffic, use the Browser to API candidate route only after choosing browser exploration as the right interface. Capture with browser-trace, generate a spec, then review/redact the report before writing any durable client. Do not treat this route as a substitute for official APIs, partner access, or terms-of-service review.

Choose the smallest durable integration surface. For a narrow stable service, an owned HTTP wrapper can be better than a large SDK when it preserves raw headers/bodies, centralized auth, retry/idempotency policy, tracing, and a small testable dependency surface. Prefer an official SDK when protocol complexity, signing, streaming, generated types, pagination, or provider support outweighs that benefit. Add Cloudflare-style joined primitives or an Eve-style agent framework only when the workflow needs their combined state, execution, observability, deployment, and operations—not because a demo compressed many products into one stack. Every route keeps a raw diagnostic/export escape and a provider migration boundary. Source: “Why we stopped using SDKs,” X Article 2077106065959989248; “Building Agents with Vercel's Eve Framework,” X Article 2069825847729508352; X 2079825908547060206, reviewed 2026-08-12

Recorded browser workflows and spatial context are input adapters, not timeless automation. Use them only on an authorized target and record browser/profile identity, account/tenant, viewport, source revision, screenshots or coordinates exposed, semantic locators, expected states, action authority, and correction path. Replay against layout/content drift, keyboard and accessibility paths, loading/error states, and one adverse fixture before promotion. Keep screenshots, form values, cookies, and recording traces under the target's privacy policy; never turn a recorded send, purchase, delete, or publish action into reusable authority. Source: X 2074973272463310905, 2077807711756968025; zk1tty/rebrowse-app@708bda96, reviewed 2026-08-12

For durable Slack capture, use scripts/sync-slack.ts: it calls only Slack's read/list/history/thread/file routes, validates the exact OAuth scope set, isolates workspaces by server-derived identity, freezes attachments, and binds edits to immutable semantic-unit revisions. Keep browser automation for user-directed visual inspection only because opening a conversation can change read state. digimata/slack 0.5.2 remains a retained alternative candidate, not the default: its JSON, ambiguity refusal, local mode-0600 credentials, bounded 429 retry, and dry-run identity controls are useful, but its source had no detected tests and its renderer is only a proposal. Never use its unsupported xoxc/cookie fallback. Source: X 2084056274451447902; digimata/slack@950569ce; Slack Conversations API, reviewed 2026-08-12

Ecosystem Pages

Tier-1 service ecosystems must have a canonical wiki/tools/<ecosystem>.md page before they are treated as promoted in active runtime indexes. That page owns the readable product-level synthesis: when to use the ecosystem, what not to confuse it with, routing, install/use guidance, failure modes, and verification path. Skills still own procedures, and source/ingest pages still own raw evidence. Source: User request, 2026-06-27; Ecosystem Capability Coverage

This is the guardrail that separates capability concepts from vendors: Browserbase Ecosystem is a browser-workflow ecosystem; Browser Testing Skills is the default agent browser route; Browser Testing Skills is the committed regression-test route; Browser Harness - Self-Healing Browser Automation is a specialist fallback. TanStack Ecosystem is a frontend data/routing library ecosystem; Cloud, Data, and Service Skills, Firebase Ecosystem, and ClickHouse Ecosystem remain backend/data ecosystems. Sentry Ecosystem, Datadog Ecosystem, and PostHog Ecosystem all sit under observability, but they answer different questions.

The cluster-level canonical pages are Identity, Session, and Routing, Browser Automation and QA, Local Web Verification, Agent Search and Retrieval, Skills, Tools, and Capability Discovery, Google Workspace Skills, Firebase Skills, Azure Skills, Cloud, Data, and Service Skills, Frontend, Design, Motion, and Accessibility, and Agent Harness, Runtime, Memory, and Evals. Use these as the readable synthesis layer before diving into individual tool pages or executable skill files.

Maintenance Rule

When adding a new tool or skill:

  1. Add the fact page (wiki/tools/ or wiki/concepts/) if the capability is durable.
  2. Add or update the executable skill only if there is a repeatable procedure.
  3. Run npm run skill-registry so executable coverage and family routing stay current.
  4. Add a resolver row in [[SKILL-RESOLVER]].
  5. Add or update this map if the change affects a capability family, tool boundary, or first-route decision.
  6. Mirror compact routing into harness projections only after the generic pages are correct, e.g. config/cursor/rules/tool-hierarchy.mdc, generated agent docs, or future Codex/Claude/OpenClaw/Hermes projection files.
  7. Run npm run ecosystem-coverage:audit when the change affects a service ecosystem.
  8. Run npm run routing-doctor, npx tsx scripts/doctor.ts --quiet, npx tsx scripts/build-index.ts, and qmd update && qmd embed.

Timeline

  • 2026-09-25 | Added sandbox provider selection while retaining the existing execution-tier and credential boundaries. Source: User request; Agent Sandboxes primary-source review, 2026-09-25

  • 2026-08-12 | Re-ran identical current Mintlify questions against the publisher's live read-only MCP and Context7. Publisher search was faster and more exact in all three observations; Context7 led with the wrong Admin MCP boundary for a docs-search question and omitted current parameter ranges. Mintlify Index returned HTTP 429 twice with a shared 1,000-request daily quota and a latest 15-hour retry window, so it remains optional discovery rather than a default. Source: X/@rohandevs; reviews/adoption-evidence/context7-vs-mintlify-documentation-2026-08-12.json]

  • 2026-08-12 | Consolidated the five-post Sentry release thread into one conditional project-CLI and release-engineering route. Stricli, generated clients, Craft, Fossilize, and binpatch now own distinct stages and enter only for measured needs; prepare, build, sign, publish, update, and rollback retain separate authority and proof. Source: X 2082947912498024751, 2082947914599469127, 2082947917237928407, 2082947920056483928, 2082947922648281278; five exact-revision repository explorations

  • 2026-08-12 | Reopened the shallowly batched X source 2075033878352494824, resolved its truncated URL to oso95/scroll-world@71cc36d3, and admitted a distinct scroll-cinematic route. The route separates pre-rendered video from runtime 3D and DOM choreography, requires position plus velocity continuity at seams, treats native 9:16 as the mobile product, and gates provider egress, spend, degraded connectors, teardown, and independent playback proof. Source: exact repository and installed skill review, 2026-08-12

  • 2026-08-12 | Added the primary-source design-tool route after a saved article was expanded into Agentation and Tweakpane repository evidence. Fieldwork owns semantic capture/artboards/versions/proof, DialKit owns motion control, Tweakpane is a narrow MIT pane, and Agentation remains a PolyForm-Shield-guarded internal/reference adapter whose fixer cannot certify its own result. Source: X 2079178687409279303; benjitaylor/agentation@8158a97; cocopon/tweakpane@50a18c3

  • 2026-08-12 | Added Toolcraft as the distinct project-scoped route for real canvas-and-controls creative applications. Ordinary frontend stays on the incumbent design system; adoption requires pinned generated source, explicit interaction/animation/artifact intent, reference mapping, and functional/browser/export proof rather than a demo reel. Source: X 2083840902779556202; pixel-point/toolcraft@4b1e415b

  • 2026-08-12 | Added guarded phone-as-server, native terminal record/replay, and HAR-derived client routes. Each preserves the useful implementation signal while making physical safety, isolation, credentials, terms, drift, redaction, and mutation authority explicit. Source: HN 49226636, 49226901, 49229026; X 2078727284865827140, 2081108843090571479

  • 2026-08-12 | Added direct-HTTP versus SDK/platform/framework selection and a privacy-, identity-, drift-, and authority-aware route for spatial and recorded browser workflows. Source: three recovered X Articles/signals and zk1tty/rebrowse-app@708bda96

  • 2026-08-12 | Deep-expanded the ten-repository X roundup 2084114545690423407. Added OpenCodeReview as an optional deterministic PR-coverage layer, ego-lite as a guarded authenticated task-space route, Instatic as a conditional visual-CMS capability and architecture reference, and the installed Text-to-CAD family plus its versioned fabrication workflow. Refreshed the existing OmniRoute, Orca, speech-to-speech, Dive into LLMs, Matt Pocock skills, and Claude Cookbooks evidence instead of creating duplicate owners. Source: frozen repository evidence at exact revisions, 2026-08-12

  • 2026-08-12 | Added the final MCP 2026-07-28 implementation route: stateless core, explicit application state, version/SDK pinning, extensions/auth as separately proven capabilities, official conformance, and legacy negotiation. Also retained Termany as a guarded high-concurrency cockpit rather than installing it as another default terminal. Source: official MCP release/schema/SDKs; thinkany-ai/termany@e121820; X 2077657071122862515, 2082167918138106256, 2079431675873038816

  • 2026-08-12 | Retained Neko, T3 Connect, and Airgorah as explicit job routes instead of filtering them out. Added disposable-room/member/clipboard/WebRTC gates for Neko, relay/device/credential/revocation gates for T3 Connect, and written-scope plus second disruptive-action approval for Airgorah. Source: X 2077863457559286004, X 2078439256230654349, X 2082277789395501263; pinned repositories

  • 2026-08-11 | Added binary-heavy version-control routing: Git remains the ordinary-repository default; Lore is a pinned, project-scoped candidate only after representative binary workload, concurrency, authority, recovery, migration, cost, and rollback proof. Source: X/@twtayaan; EpicGames/lore@v0.8.6; official FAQ/system design/roadmap, reviewed 2026-08-11

  • 2026-08-11 | Promoted Wayfinder as the distinct route for multi-session decision discovery; clear plans remain on to-prd/to-issues, and external trackers or parallel research retain their ordinary authority gates. Source: X/@MengTo paired-library replay; mattpocock/skills exact source

  • 2026-08-11 | Added Box by ASCII as a guarded managed persistent full-VM route rather than a new page family or default dependency. Admission requires a clean no-env environment, external account key, tenant/budget/lifecycle receipts, separate public-access approval, reproducible installation, and resolution of the missing privacy policy and third-party-access terms. Source: X 2076651776527224913; Box docs/API/SDK/terms/media replay

  • 2026-08-11 | Added Coolify as the explicit recommended self-hosted-PaaS route beside the managed-PaaS route. Observation uses team-scoped read MCP; lifecycle/configuration mutations require separate deploy/write authority through the official interfaces; production admission requires owned operations, backup/restore, canary, health, rollback, and Docker-escape proof. Source: X/@sentient_agency; current Coolify repository, CLI, OpenAPI, MCP, and operations docs; capture

  • 2026-08-11 | Added evidence-gated routes for query-embedding adaptation, high-stakes mechanical decision control, and paired-objective guardrail-policy optimization. Santander references remain conditional patterns; no package became a default install. Source: SantanderAI exact-revision replay and review

  • 2026-08-11 | Added runtime memory as an explicit capability family. Kevin-Wiki/qmd remains canonical; Hindsight is retained as the selected candidate/default only after pinned install, independent local/cloud path declaration, bank/provenance policy, doctor, leakage and usefulness comparison, and fresh-target replay. Recall and reflect now route as distinct cost/behavior surfaces, with Honcho kept in the separate profile lane. Source: X/@itsharmanjot; vectorize-io/hindsight@d7c33fde

  • 2026-08-11 | Added a guarded real-time collaborative-coding route. Jam is retained and discoverable for disposable or deliberately shared work only after participant/repository, BYOK secret, egress/spend, revocation/deletion, and public-preview boundaries; the existing governed session remains first for private or production work. Source: X 2070242834985431293; Jam UI/privacy/terms/client schema and browser replay

  • 2026-08-11 | Added personal data-broker removal as a permission-first family: official California DROP before third-party automation, assisted field-level canaries for gaps, and external filing/removal proof. Unbroker is retained but gated by current source/test, mixed-license, encryption, processor, browser/email, recipe, and live-quality receipts. Source: Unbroker complete replay, current implementation, CalPrivacy DROP, and Browserbase primary docs, 2026-08-11

  • 2026-08-11 | Added authorized protected-page retrieval as a separate guarded family: permission and least-evasive routes first, measured failure classification before tool choice, and Camofox only under target-specific security and quality proof. Source: Camofox Browser full replay and current-source receipt, 2026-08-11

  • 2026-08-11 | Added rendered-layout retrieval as an explicit hybrid lane: text/DOM baseline first, PixelRAG when visual structure changes answerability, and mandatory source/render identity, region/link grounding, disagreement, privacy/rights/moderation, reader, cost, and bounded-evaluation gates. Source: PixelRAG current repository, paper, evaluation package, thread, and video replay, 2026-08-11

  • 2026-08-11 | Added the executable macOS discovery route: a tested Awesome-Mac snapshot/diff feeds candidates into official-source review, trial, doctor, and workflow binding. All 1,257 rows are retained without turning the catalog into an install manifest or inheriting its conflicting license metadata. Source: X/@XAMTO_AI; jaywcjlove/awesome-mac@bab7c005

  • 2026-08-11 | Added the agent execution substrate ladder: named capability/Code Mode → virtual OS → full worker, with explicit hybrid escalation and proof. Secure Exec remains a specialized agentOS compatibility/API surface rather than a duplicate route; blanket vendor security and density claims remain held. Source: X 2071672946859688352; current agentOS and Secure Exec primary sources

  • 2026-08-11 | Split the five-link coder-app bookmark into job-shaped routes: HeyForm for deliberately operated conversational forms, Checkmate for self-hosted availability/hardware monitoring, Sevalla for measured low-operations PaaS needs, DocMD for new standalone agent-readable docs, and VARCHIVE as frontend discovery evidence. None became a global install or a replacement for Kevin Wiki. Source: X/@csaba_kissi; source snapshots]

  • 2026-08-11 | Replayed the five-app coder roundup as five distinct capabilities: API Finder is a candidate directory below provider-owned docs; TraceDR is an auxiliary Domain Rating trend signal below first-party SEO evidence; Orca remains an optional fleet client; CSS Loaders is a bounded semantic-wait treatment; and React Bits remains a restricted-license component source. Source: X/@csaba_kissi; Source snapshots

  • 2026-08-10 | Added OpenCLI beneath agent-browser as a guarded repeated-adapter route, with separate-profile/project isolation, exact adapter/permission/dependency review, and explicit mutation authority. It remains uninstalled globally. Source: X/@csaba_kissi; Source: OpenCLI repository review, 2026-08-10

  • 2026-08-10 | Replayed a ten-repository open-source roundup without treating it as one recommendation. Added bounded routes for Open Notebook and Odysseus research workspaces, PhotoGIMP, FreeDomain, WebToApp, and ReClip; existing Recordly, Stirling-PDF, HyperFrames, and Excalidraw routes were reinforced. Source: X/@DivyanshT91162; current first-party repositories

  • 2026-08-10 | Added Kimi K3 as a guarded multimodal, million-token, open-weight model candidate. Same-harness local receipts, provider/privacy, custom-license, cost, and hardware gates now precede promotion; mixed-harness vendor benchmark images do not establish a default. Source: X/@Kimi_Moonshot; Kimi K3 repository, report, API, and license

  • 2026-08-10 | Routed the six-repository builder roundup atomically: Firecrawl for structured web extraction/crawl escalation, Agent Reach for multi-platform channels, Recordly for human-operated capture, Website Cloner for authorized structural reconstruction, Open Generative AI as a guarded local/UI reference, and SkillClaw as a session-to-skill proposer under evidence-gated promotion. No external skill bundle was globally installed. Source: X/@aashatwt; frozen first-party repositories and capture manifest

  • 2026-08-10 | Added the research-experiment route: the evaluator and provenance own the claim; ML Intern is a conditional HF-native executor whose local-default, headless-YOLO, cost, and trace-egress boundaries must be contained before sensitive use. Source: X/@akseljoonas; Hugging Face ml-intern, reviewed 2026-08-10

  • 2026-08-10 | Added a guarded structural-retrieval route: SQL first, pgGraph for repeated bounded Postgres traversal, and Polygres only for measured hybrid needs; qmd remains the wiki default and pgGraph topology is excluded where RLS/tenant behavior is unproven. Source: X/@daleverett X Article; pgGraph/pgContext/Polygres primary sources, reviewed 2026-08-10

  • 2026-08-10 | Converted the cross-department skills bookmark into explicit routing: Context7 is a current-docs escalation after local and official versioned docs; UI UX Pro Max remains a specialized external design source; marketing and social collections are selective source catalogs rather than bulk installs; and Anthropic's Finance, Small Business, and Legal plugins are Claude Cowork routes with connector, approval, and professional-review boundaries—not portable cross-agent capabilities. Source: X/@cyrilXBT, 2026-07-16; Source: current GitHub and Anthropic plugin primary sources, reviewed 2026-08-10

  • 2026-08-10 | Added three bounded exceptions from the follow-on GitHub frontier review: OmniRoute for locally governed multi-provider gateway experiments after security and provider-terms review; AstrBot for a full self-hosted multi-channel chatbot runtime after AGPL/plugin/sandbox review; and KTransformers for measured large-MoE CPU/GPU inference or fine-tuning rather than ordinary local serving. Source: GitHub diegosouzapw/OmniRoute, AstrBotDevs/AstrBot, and kvcache-ai/ktransformers, captured 2026-08-10

  • 2026-08-10 | Reviewed the duplicated GitHub-frontier roundups and refined three routes: Graphify remains first with Code Review Graph as a conditional local MCP alternative; Hallmark complements rather than duplicates Impeccable; OfficeCLI may back the existing document/spreadsheet/slide skills when its render loop is useful, but its auto-installer must not mutate every harness blindly. Source: GitHub tirth8205/code-review-graph, Nutlope/hallmark, and iOfficeAI/OfficeCLI, captured 2026-08-10

  • 2026-07-14 | Added technology-vitality, chart-engine/treatment, and Impeccable design-language routes; refreshed ACP and Hermes memory boundaries from their official docs. Source: User request, 2026-07-14

  • 2026-07-02 | Added Graphify as the safe repo-topology sidecar route, explicitly below qmd for wiki knowledge and above raw file sweeps for structural graph/path/affected-node questions. Source: User request, 2026-07-02

  • 2026-06-30 | Added Exa Agent as the async structured research/enrichment route between live-source research and generic WebSearch; it is for list-building, row enrichment, and schema-backed research, not for Kevin wiki recall or low-latency lookup. Source: X/@ExaAILabs, 2026-06-16

  • 2026-06-30 | Promoted LiteParse into the document/media extraction route as the fast local PDF/layout parser with an imported executable skill; MarkItDown - Universal File-to-Markdown Converter remains the broad ordinary file converter and MinerU/OmniParse remain escalation routes. Source: X/@jerryjliu0, 2026-05-27; Source: GitHub/npm/PyPI, 2026-06-30

  • 2026-06-30 | Surfaced Photon Spectrum as the messaging-channel agent UI route and Its Hover / Super Hover as deliberate animated hover/icon affordance references so routing-doctor can discover their tool pages. Source: routing-doctor, 2026-06-30

  • 2026-06-30 | Added Browser to API as the candidate route for undocumented website API discovery, with explicit redaction and API-contract caveats. Source: X/@derekmeegan, 2026-05-13

  • 2026-06-29 | Made this page the explicit harness-agnostic source of truth for tool routing; Cursor tool-hierarchy.mdc is a projection that mirrors this map and Skill Resolver. Source: User correction, 2026-06-29

  • 2026-06-29 | Added Loopy and agent self-improvement evals as capability families so repeatable loop design and evaluator hardening route before generic skill creation. Source: User request, 2026-06-29

  • 2026-06-27 | Linked ecosystem capability coverage as the pre-consolidation check for service ecosystem promotion across skills.sh, MCP/CLI, wiki pages, routing, packs, and verification. Source: User request, 2026-06-27; scripts/audit-ecosystem-coverage.ts

  • 2026-06-27 | Replaced skill/tool mirror guidance with the executable skill + generated registry + family synthesis model. Source: User request, 2026-06-27; scripts/generate-skill-registry.ts

  • 2026-06-18 | Linked Agent Capability Registry as the discovery layer that feeds this routing map. Source: whole-wiki graph audit, 2026-06-18

  • 2026-06-18 | Added open conditions and the interface cost ladder so future agents know when to use the map and when to stay with a specific skill. Source: User request, 2026-06-18

  • 2026-06-17 | Created after Kevin clarified that duplicated skill/tool entries are acceptable when they represent different layers; the needed cleanup is capability hierarchy and routing, not blind deduplication. Source: User, 2026-06-17

75 pages link here

Active Stack (What to Actually Use)MetaAgent Browser -- Browser Automation for AgentsToolsAgent Capability RegistryConceptsAgent Company OSArchitectureAgent Config Write ChecklistHEARTBEATAgent Docs FilingFiling Decision TreeAgent Docs System MapAgent DocsAgent Environment Setup ArchitectureArchitectureAgent ExpectationsUSERAgent HarnessConceptsAgent Operating SystemArchitectureAgent Operations HubMetaAgent Runtime ArchitectureArchitectureAgent Search and RetrievalConceptsAgent Security ModelConceptsAgent SoulSOULAgent Soul HubSOULAgent-Native CLIsConceptsBrowser Automation and QAConceptsCapability Harvest PatternMetaCapture Ingest ProtocolMetaCapture Ingest QuickrefAgent DocsCapture Ingest WorkflowWorkflowsCoder Apps BundleToolsCurrent WorkUSERCursor Rules OverviewAgent DocsEcosystem Capability CoverageMetaego-liteToolsExa AgentToolsFiling Decision TreeFiling Decision TreeFMHYToolsFrontend, Design, Motion, and AccessibilityConceptsFull Corpus Workflow and Capability ProgramArchitectureGraphifyToolsGraphify Sidecar WorkflowWorkflowsHow to Update the WikiMetaHugging Face Speech to SpeechToolsIdentity, Session, and RoutingConceptsImpeccableToolsInstaticToolsIs This Tech Dead?ToolsKevin Agent OS Whiteboard — July 2026ArchitectureLocal Search CLIProjectsLunoraToolsMCP as the Integration StandardConceptsMCP RegistryToolsMemory Routing ProtocolMetaNative SDK (formerly Zero Native)ToolsOpenCodeReviewToolsOpenRouterToolsOpenRouter FusionToolsOperational HeartbeatHEARTBEATOrcaRouterToolsPaperclipToolsPDF OperationsToolsProduct Site Reference GalleryDesignQuality GatesMetaRealtime Voice Agent WorkflowWorkflowsResearch Experiment WorkflowWorkflowsResolver vs Skill ResolverFiling Decision TreeShopifyToolsSkill ResolverSkill ResolverSkill Resolver FilingFiling Decision TreeSkills FirstSOULSkills, Tools, and Capability DiscoveryConceptsSlash Command IndexMetaTechnical StackUSERText to CAD Skill FamilyToolsTool Hierarchy CompactAgent DocsTool Routing DisciplineSOULTools FilingFiling Decision TreeUser ProfileUSERWiki Capsules And Agent Doc PacksMetaWiki Concept Coverage MapMetaX Bookmarks: AI Agents & Tools (Jan 2025 – Jun 2026)Tools