Agent Control Plane Architecture

An agent control plane manages fleets of agents and workers: provisioning, identity, credentials, routing, schedules, policies, traces, approvals, usage, and recovery.

Boundary

An agent control plane is not the agent. It is the system that makes agents operable at fleet scale.

Plane Owns Example
Data plane The actual running agent, tools, sandbox, process, browser, filesystem, and model calls A Claude Code process in a VM, a Hermes worker, a LangGraph run
Control plane Desired state, provisioning, routing, credentials, schedules, policies, observation, approvals, lifecycle Agent Machines, a cloud agent console, a managed workflow dashboard

The analogy is Service Mesh and Service Discovery Architecture and Kubernetes: the data plane does the work, while the control plane decides what should exist, where it runs, what it may access, and how operators see it.

Agent Company OS gives the same control-plane idea an organization model: department leads own KRs, memory-skill overlays keep workers scoped, and the runtime command plane owns wakeups, cron, permissions, model routing, and proof ledger state. Source: X/@DerekNee image, 2026-06-25

Core components

Component Responsibility
Identity and tenant model Users, orgs, service accounts, agent identities, access boundaries
Capability registry Approved runtimes, skills, MCP servers, CLIs, tools, model providers, and presets
Provisioner Creates and destroys sandboxes, persistent machines, browsers, volumes, and runtimes
Credential broker Stores and injects model keys, service tokens, OAuth grants, and scoped secrets
Router Chooses runtime, model, substrate, region, and tool path for a workload
Scheduler Cron, timers, wakeups, queued runs, backoff, and retry policies
Run/session store Durable record of runs, machine state, artifacts, traces, approvals, and logs
Policy engine Permissions, approval gates, spend limits, network rules, tool allowlists
Observation surface Live status, logs, traces, screenshots, costs, health, and task outcomes
Operator interface Dashboard, CLI, MCP/API, webhooks, Slack/mobile approvals
Billing and quotas Usage attribution across model calls, machines, storage, and tool providers

Control flow

IRL comparison

System Control-plane lesson
Kubernetes Desired state, reconciliation, scheduler, API server, and controllers make many containers operable. Agents need a similar management layer when they become long-lived workers.
Vercel / cloud PaaS Developers pay for deploy, routing, logs, env vars, preview URLs, rollback, and domain wiring, not just a Node process. Agents need the same productization.
LangSmith / managed LangGraph Tracing, evaluation, deployment, and prompt/version management become the operational layer around LangGraph runs.
Agent Machines The product thesis is a neutral control plane for persistent skilled workers across runtimes, substrates, skills, and model routes.
Paperclip A current MIT implementation of company/goal/task/heartbeat/budget/approval/workspace/secrets/activity control across heterogeneous agent adapters.
Enterprise MCP/governance tools Capability discovery is not enough; orgs need approval, inventory, audit, and revocation. See Agent Capability Registry.

Field Evidence From Paperclip

Paperclip makes several abstract requirements testable. A task can carry goal and parent ancestry, blockers, a single assignee, an atomic checkout/execution lock, workspace identity, review participants, cost events, and a run ID on mutations. Heartbeats are database-backed wakes with coalescing, budget checks, session continuation, secret/skill injection, adapter execution, cancellation, and orphan recovery. Routines turn cron/webhook/API triggers into tracked issues instead of running invisible scheduled prompts. Source: paperclipai/paperclip README, product definition, agent skill, and server source at 66575fe, 2026-08-11

Those features also expose the real adoption gates:

  • A control plane only observes what its adapters report. “Every tool call” is not a safe universal claim for opaque external runtimes.
  • Local no-login mode is loopback-only; private/public network use requires the authenticated deployment model and recovery-aware identity/bootstrap flows.
  • Secret encryption needs key + database + file backup ownership, and company scoping still needs adversarial isolation tests before hostile multi-tenancy.
  • Anonymous product telemetry is enabled by default even though OpenTelemetry tracing is opt-in; privacy reviews must distinguish the two.
  • Company skill mutations are open unless an explicit policy restricts them; core safety checks do not substitute for least-privilege operating policy.
  • Current master had two failing workspace-busy scheduled-retry tests in one CI shard at review time. A production pilot must pin a green release/revision and explicitly exercise collision, retry, and recovery semantics.

Kevin-Wiki can project approved workflows and goals into a control plane, but the control plane remains replaceable runtime state. Source evidence, canonical objects, accepted instructions, and proof-governed learning stay in the brain.

2026 state of the art

  • Neutral routing. Model, runtime, and substrate are separate axes. Agent Machines frames this as OpenRouter for agents and containers.
  • Persistent workers. The control plane needs pause, wake, fork, snapshot, volume, and resume semantics, not only "start a job."
  • Recursive operation. A human is one caller, but a head agent may also provision workers through MCP/API and supervise them.
  • Governed capabilities. MCP, skills, CLIs, and remote agents must pass through a registry and policy layer before use.
  • Trajectory-aware observability. Logs are insufficient. Operators need model/tool/handoff/approval traces and artifacts to evaluate behavior.
  • One-click cancel and rollback. Long-running autonomy without interruption paths is a liability.

Invariants

  • The control plane is the source of desired state; workers report observed state.
  • A worker can disappear without erasing the run record.
  • Credentials are injected by policy, not pasted into prompts.
  • Routing decisions are recorded so failures can be reproduced.
  • Every destructive or costly capability has an approval or budget boundary.
  • Operators can answer: who ran what, where, with which tools, at what cost, and with what outcome?
  • Department leads can dispatch work, but worker output is not complete until it returns proof artifacts that update the workspace brain or memory/skill system.

Control-plane outage contract

Static stability keeps existing data-plane work available without requiring incident-time provisioning. PlanetScale describes separating query serving from database management; AWS describes reserving enough capacity to tolerate the chosen availability-zone failure before it occurs. These are design principles, not a guarantee that every dependency or regional failure is isolated. Source: PlanetScale, Max Englander, 2025-07-03; AWS Builders’ Library, Becky Weiss and Mike Furr; reviewed 2026-09-10

For agent workers, the proposed architectural translation is to define which already-authorized work can continue with locally available configuration and reserved capacity, and which operations require a healthy control plane. Continuing work must preserve locally enforceable scope, budget, expiry, audit, and cancellation requirements. If those requirements cannot be enforced, pause the affected work. Cached state cannot grant a new approval, extend an expired grant, or bypass a required revocation check. This is a design inference; deployed worker behavior requires separate verification.

Verify the contract by interrupting management dependencies and a worker failure domain: measure continuing work, rejected operations, retained evidence, and recovery capacity. Record the tested failure scope and authority limits rather than claiming universal outage immunity. See Resilience Patterns (Circuit Breaker, Bulkhead, Retry) for call-site defenses.

Kevin-stack implication

Any serious agent product in Kevin's ecosystem should identify whether it is building:

  • a harness,
  • a runtime,
  • a control plane,
  • a capability registry,
  • or a vertical preset on top.

Confusing those layers is how products become a pile of buttons around a terminal. The control plane is the company when the user is paying to avoid rebuilding deployment, observation, permissions, and routing.

Architecture Position

Axis Value
Family Agent runtime and control plane
Boundary owned Fleet operation: provisioning, identity, credentials, scheduling, observation, policy, billing.
Read with Agent Runtime Architecture, Agent Harness Architecture, Agent Machines
Use this page when designing a platform that manages many agents/workers

Timeline

  • 2026-08-11 | Grounded the control-plane model in current Paperclip source: goal ancestry, atomic checkout, bounded heartbeats, routines, budgets, review/approval, workspaces, secrets, activity, and heterogeneous adapters; added telemetry, policy, isolation, retry-CI, and canonical-memory adoption gates. Source: X/@NickSpisak_; GitHub paperclipai/paperclip HEAD 66575fe

  • 2026-07-01 | Architecture category refresh added this page to the Agent runtime and control plane family, linked it to Architecture System Map, and kept it standalone because it owns this boundary: Fleet operation: provisioning, identity, credentials, scheduling, observation, policy, billing. Source: User request, 2026-07-01

  • 2026-06-25 | Added Agent Company OS as an org-chart framing for the control plane: department leads, memory overlays, worker pools, and proof ledger state. Source: X/@DerekNee, 2026-06-25

  • 2026-06-18 | Created as the missing architecture page behind Agent Machines, browser agent console, and recursive agent orchestration. Source: whole-wiki graph audit, 2026-06-18