Agent Company OS

Agent Company OS is a multi-agent operating model: scoped workspace brain, memory/skill overlay, runtime command plane, department lead agents, specialized workers, and a proof loop that writes work back into memory.

Agent Company OS blueprint

Derek Nee's "Matrix Operating Blueprint" is a concise architecture diagram for autonomous work. The core claim is that company-scale agent work cannot run through one giant omniscient agent. The system needs boundaries, operating rules, accountable leads, scoped workers, and proof artifacts. Source: X/@DerekNee, 2026-06-25

Architecture

The diagram has five load-bearing layers:

Layer Job
Company context Source material enters as assets, rules, past experience, and examples of taste.
Workspace brain Holds the boundary for one workspace: objectives, departments, messages, evidence, and rules.
Memory + skill system Stores what agents read before acting: long-term memory, skills, runbooks, examples, project constraints, and reusable taste.
Runtime command plane Coordinates wakeups, cron, messages, objective state, proof ledger, permissions, and model/runtime routing.
Department agents and workers Long-running department leads own accountability, then dispatch work to the right execution seat: Codex, Claude Code, native workers, or browser/computer workers.

The final layer is the proof loop. Artifacts and traces satisfy criteria, update objective state, write check-ins, and improve the memory/skill system. Source: X/@DerekNee image, 2026-06-25

Concrete Control Plane: Paperclip

Paperclip turns much of this diagram into an inspectable implementation. Its current source models companies, goals, projects, issues, parent/blocker graphs, humans and agents, reporting lines, heartbeats, routines, budgets, approvals, execution workspaces, cost events, secrets, and activity. Claude Code, Codex, Cursor, Gemini, OpenCode, Pi, process/HTTP, OpenClaw, Hermes, and plugins remain execution adapters; the control plane owns coordination and observed state. Source: paperclipai/paperclip HEAD 66575fe, reviewed 2026-08-11

The implementation sharpens this page in four ways:

  1. Goal alignment needs a traversable task graph. A slogan or mission prompt is insufficient; every issue needs a parent/goal path, explicit blockers, accountable assignee, and completion evidence.
  2. A heartbeat is a bounded execution lease. It wakes with identity, issue, budget, workspace, and authority; it exits with durable status, artifacts, costs, and a next wake path. Narrating “waiting” is not resumability.
  3. Governance must be enforced in state. Atomic checkout, run attribution, typed approval/review stages, hard budgets, cancel/pause, scoped secrets, and recovery paths matter more than anthropomorphic titles.
  4. The org database is runtime state, not the brain. Paperclip can retain tasks and sessions, but accepted learning still writes through Kevin-Wiki's source/object/review contracts. Company skills and task comments do not silently become canonical instructions.

Paperclip's own boundary is equally useful: it says a single-agent task does not need a company control plane. Kevin's default remains one named workflow in one appropriate harness. Add an organization layer only when multiple persistent workers need shared goal state, schedules, budget attribution, approval, recovery, and operator visibility. The reviewed master revision also has one failed CI shard around workspace-busy retries, so the tool remains pilot-gated rather than assumed production proof. Source: Paperclip README, product definition, deployment docs, and GitHub checks, 2026-08-11

Design Implications For Kevin's Wiki

This pattern maps directly onto Kevin's existing operating system:

The operational upgrade is to keep agents narrow without making them blind. A worker should get the relevant memory skill overlay, not the whole wiki or every tool. A lead agent should own routing, approvals, and proof, not every low-level action.

Greg Isenberg's 2026 diagrams add a useful operator thesis: make the company legible, give agents repetitive but meaningful work, and remove avoidable human handoffs. They are not an enforcement model. “Humans do strategy; agents do execution,” “one shared context layer,” and “continuous learning” all need stricter authority, data, and proof contracts before they are safe operating rules. Source: X/@gregisenberg, 2026-06-27

Compile The Operator Thesis Into Policy

The post's “shared context layer” is a useful logical picture, but it must not become one prompt, one vector database, or one permission domain. Company legibility is a typed context fabric whose planes remain separately owned:

Plane Contains Required boundary
Source archive Messages, documents, recordings, external pages, and raw data Immutable revisions, provenance, visibility, retention, and deletion state
Compiled knowledge Accepted facts, decisions, examples, runbooks, and design rules Canonical owner, supporting and contradicting evidence, freshness, and review status
Operational state Goals, tasks, customers, workflow state, budgets, schedules, and incidents Current revision, accountable owner, lifecycle, idempotency, and recovery
User/team memory Preferences, relationships, corrections, and interaction history Subject, purpose, consent, visibility, expiry, and revocation
Policy and authority Roles, permissions, approval rules, data boundaries, and risk classes Deny-by-default evaluation outside the worker and an auditable grant/decision
Secrets Credentials, tokens, signing material, and private endpoints Brokered just-in-time access; never copied into general retrieval or durable prompts
Decision and proof ledger Proposals, approvals, traces, artifacts, outcomes, and evaluator results Actor, target digest, evidence, cost, rollback, and acceptance state

Every retrieved item therefore needs an owner, source revision, visibility, authority class, freshness or validity window, and revocation path. The runtime compiles the smallest task-specific context bundle; a worker does not receive the whole company because it can search the company. Wiki Brain Operating Model implements this split for Kevin-Wiki through separate source, compiled, user, run-state, operating, and proof surfaces.

Human and agent authority is decided per action

The org chart describes comparative strengths, not blanket authority. Strategy can be researched and drafted by an agent; repetitive execution can still be too consequential to automate.

Work property Default operating mode
Cheap, reversible, well-specified, and deterministically checkable Agent may execute inside a standing envelope and return proof.
Novel, ambiguous, taste-bearing, or goal-changing Agent investigates and proposes options; the accountable human chooses the direction.
External, irreversible, expensive, privacy-bearing, or reputation-sensitive Mechanical policy gate first, then explicit human approval of the exact action digest.
Legal, medical, financial, employment, safety, or regulated judgment Agent assists within a narrow role; the qualified accountable decision-maker retains the decision and required review.
Repeated behavior with stable inputs and measured outcomes Candidate for broader automation only after eval, failure, rollback, and escalation evidence.

Use Human-in-the-Loop Control for the exact pause/approve/edit/reject/resume contract. Functional labels such as “legal agent” or “finance agent” name a capability lane, not an employee-equivalent license to act.

Opportunity ranking is a gated scorecard

Repetition plus workflow complexity finds interesting markets, but it does not establish that a workflow is safe, feasible, or economic. Score each candidate on the following before funding an agent loop:

  1. Volume and repetition — enough recurring work and examples to justify a system.
  2. Process and data readiness — inputs, rules, exceptions, and rights are discoverable and usable.
  3. Outcome observability — success, failure, and partial success can be measured without self-report.
  4. Verifiability — deterministic checks, comparative evaluation, or qualified review can catch material errors.
  5. Reversibility and blast radius — mistakes can be contained, cancelled, or compensated.
  6. Exception entropy — the tail of novel cases is known enough to route rather than guess.
  7. Authority and stakes — the system can obtain legitimate authority without collapsing separation of duties.
  8. Integration readiness — reliable APIs, identity, state, and failure recovery exist at the workflow boundary.
  9. Unit economics and latency — successful outcomes remain worthwhile after model, runtime, retry, support, and human-review cost.
  10. Fallback quality — an owner, escalation path, dead-letter state, and graceful manual route exist.

Do not let a high aggregate score override a red line. If success cannot be observed, authority cannot be established, or failure cannot be contained, start with retrieval, drafting, or recommendation rather than autonomous mutation. Agent Unit Economics owns the cost-per-successful-outcome test.

Compress the maze, not the control boundary

The durable “after” workflow is:

accepted request
  -> bounded context retrieval
  -> policy and authority check
  -> propose or execute within the granted envelope
  -> deterministic/domain validation
  -> typed human gate when required
  -> commit, respond, or publish
  -> proof receipt
  -> learning proposal

Remove inbox routing, duplicate lookup, transcription, and status handoffs. Preserve required approvals, separation of duties, counterparty rights, and independent verification. Every mutating path still needs target revision, idempotency, stale-event rejection, timeout, cancel, retry, escalation, dead-letter recovery, and rollback. Faster passage through a missing control is not workflow improvement.

Continuous learning is an admission pipeline

An outcome does not silently teach the company. The safe loop is:

run evidence -> learning proposal -> independent verification
  -> typed acceptance -> versioned writeback -> regression and rollback proof

Raw conversations, worker memories, task comments, and successful-looking outputs remain evidence until this loop accepts them into a canonical owner. The proposer cannot be the only verifier, and a metric improvement cannot erase privacy, authority, reliability, or policy regressions.

Legibility is necessary, not a moat by itself

A readable company is easier for its own agents to operate, but also easier to copy or migrate. Defensibility can emerge from lawful proprietary feedback, trusted distribution and integrations, domain-specific evals, operational reliability, data rights, switching cost, and a learning loop that admits the right changes. Legibility is the prerequisite that makes these assets usable; it is not sufficient evidence of durable advantage.

The source remains valuable as a high-engagement operator thesis. LCA's public site corroborates that the firm offers AI-native product work, custom agents, workflow automation, and team enablement, but its public marketing and case studies do not measure the organizational outcome or moat claimed in the post. Treat those claims as hypotheses to evaluate, not reported performance. Source: LCA homepage, captured 2026-08-11; LCA AI Acceleration, captured 2026-08-11

Giga Scout is the hosted-product version of the same control problem. Its public model starts from a business KPI, learns from real customer conversations, files fixes as policy/tooling/knowledge changes, tests safe changes on traffic slices, and escalates risky changes to humans. That makes the metric and rollout policy first-class parts of the agent company OS rather than after-the-fact reporting. Source: Giga Scout, 2026-07-02

Each diagram still contributes a compact test: the org chart asks who retains accountability; the stack asks whether context, policy, review, and learning are connected; the opportunity map asks where repeated complexity creates value; and the before/after flow asks which handoffs can disappear. The compiled rules above supply the missing authority, evidence, failure, and economic gates. Source: X/@gregisenberg, 2026-06-27; Source: local review of four image artifacts, 2026-08-11

Workspace Pod Rule

Each workspace should be its own pod: separate brain, memory, tools, workflows, approvals, and proof ledger. Fork the pattern, not raw context. This supports Kevin's existing kevin-wiki, Dedalus, Agent Machines, Loop, and project-specific agent-docs meshes: each project can share the same operating model while keeping local constraints and evidence scoped.

Anti-Patterns

  • One agent with every file, every tool, and no accountable boundary.
  • A command room with no proof ledger.
  • Workers that can spawn more workers without recorded objective, budget, or cancellation state.
  • Memory that only accumulates context and never updates skills, rules, or runbooks.
  • Department-like pages or skills that nobody routes through.
  • An org chart wrapped around one task or one agent.
  • Treating a control-plane task database, agent chat, or mutable company skill as canonical memory.
  • Treating “shared context” as one prompt, vector store, or permission domain.
  • Letting run outputs, engagement, or a single KPI silently rewrite memory, policy, or skills.
  • Selecting work from repetition and complexity alone while ignoring authority, verifiability, exception entropy, reversibility, and unit economics.
  • Removing a required approval or separation-of-duties boundary in the name of workflow compression.
  • Treating company legibility as sufficient proof of defensibility or operating performance.
  • Scheduled heartbeats without no-work suppression, per-goal budgets, cancel paths, and a successful-outcome metric.
  • Calling a self-hosted control plane private while default telemetry, public exposure, open skill policy, secret recovery, or backups remain unreviewed.

Timeline

  • 2026-08-11 | Replayed Greg Isenberg's post and all four original diagrams under the source-review contract; replaced the slogans with typed context planes, action-level authority, a ten-part opportunity gate, control-preserving workflow compression, proof-gated learning, and an evidence-qualified moat claim. Preserved LCA's official product/team assertions as corroboration of activity, not outcome proof. Source: X 2070918939526205494; local artifact review; LCA official site

  • 2026-08-11 | Added Paperclip as a concrete, source-pinned company control plane and tightened the model around traversable goal ancestry, bounded heartbeat leases, state-enforced governance, canonical-memory separation, single-agent avoidance, and pilot proof. Source: X/@NickSpisak_; GitHub paperclipai/paperclip HEAD 66575fe

  • 2026-07-03 | Deep-reviewed Greg Isenberg's four local diagrams and added the operator checklist: humans own judgment, shared context is the real stack, opportunities need repetition plus complexity, and maze-like workflows compress into policy-gated agent loops. Source: X/@gregisenberg, 2026-06-27; Source: local artifact review, 2026-07-03

  • 2026-07-02 | Added Giga Scout as a concrete KPI-governed agent-company product example: define the metric, learn from conversations, file fixes, test on slices, escalate risky changes, and write accepted changes back into policy or knowledge. Source: Giga Scout, 2026-07-02

  • 2026-06-29 | Added Greg Isenberg's operator-facing agent-company diagrams: humans move to strategy/taste/judgment while agents execute bounded loops, making goal, proof, metric, and escalation state load-bearing. Source: X/@gregisenberg, 2026-06-27

  • 2026-06-25 | Page created from Derek Nee's Agent Company OS blueprint and mapped into Kevin's wiki/skills/automation architecture. Source: X/@DerekNee, 2026-06-25