Agent Sandboxes

Sandboxes started as safe environments for running untrusted code. As agents become persistent and stateful, the sandbox is evolving into the operating environment agents actually work inside. Source: X/@crystalxtang, 2026-05-27

Three workload shapes

Sandbox use cases split along two axes, how long they run and how much state they hold, which collapse into three dominant shapes. Source: X/@crystalxtang, 2026-05-27

  • Simple code execution: short-lived environments that spin up, run code, and disappear. The default for most LLM tool calls today.
  • Stateful environments: longer-lived, dev-like sessions with persistent filesystems and multi-step workflows that last hours.
  • Compute-heavy workloads: GPU-bound, parallelizable, cost-intensive jobs (small-model training, inference, batch, CI).

Provider selection: September 2026 review

Choose by workload and existing platform before comparing speed claims. This is a documentation-backed shortlist, not a benchmark ranking or a decision to adopt every provider.

Provider Relevant workload or boundary Primary evidence
Daytona Daytona Stateful agent workspaces with filesystem, Git, process, and language-server APIs.
Blaxel Blaxel Managed virtual-machine sandboxes with automatic standby and resume.
Cloudflare Sandbox Sdk Cloudflare Sandbox SDK Workers-controlled Linux execution through Cloudflare Containers.
Vercel Sandbox Vercel Sandbox Vercel's isolated execution service with persistent sandboxes, snapshots, and SDK/CLI access.
Docker Sandboxes Docker Sandboxes The sbx interface for local coding-agent isolation and Docker-managed cloud sandboxes.
Cube Sandbox CubeSandbox An open-source RustVMM/KVM sandbox service with an E2B-compatible API.
Smol Machines Smol Machines A local microVM runtime, hosted cloud service, and optional application SDK.
Aws Fargate AWS Fargate Managed container execution for ECS workloads with per-task isolation.
Runloop Devboxes Runloop Devboxes Managed virtual workstations for repository, browser, and command execution.
Declaw Declaw A sandbox platform combining Firecracker execution with external policy and credential mediation.
Hopx HopX Cloud execution environments with file, process, template, and desktop-automation interfaces.
Fly Sprites Fly.Io Sprites Persistent Linux environments with filesystem checkpoints and restore.
Codesandbox CodeSandbox Template-driven development sandboxes documented by Together AI.
Modal Labs Modal Sandbox execution alongside Python/AI compute workloads.
Compute Sdk ComputeSDK A common TypeScript interface and benchmark harness across providers.

The recovered July snapshot grouped providers into hosted execution, local/self-hosted environments, and cloud substrates. Keep that distinction, but avoid fixed provider tiers: Docker now spans local and cloud, Daytona has multiple isolation classes, and persistence is not a single capability. For Docker, moving copies a filesystem into a separate instance without carrying managed secrets, local policy, or running processes. Sprites checkpoints similarly restore disk rather than process execution. Source: Docker, Daytona, Sprites, checked 2026-09-25

Ownership and proof

Keep Kevin's rules, skills, source history, approvals, and proof outside the rented execution service. Persist portable workspace artifacts and explicit lifecycle receipts. Operate a self-hosted fleet only when data-plane ownership is itself required and funded. Vendor documentation is useful evidence of the offered interface, but vendors benefit from adoption; it cannot prove local latency, isolation, recoverability, or cost. The existing provider-fixture contract below remains the adoption gate. Source: Build, Don't Buy; Agent Security Model; current provider documentation, reviewed 2026-09-25

Package snapshot

Registry tags checked 2026-09-25; these are research observations, not installed versions or upgrade instructions.

Package npm latest
@daytona/sdk 0.218.0
@blaxel/core 0.3.23
@cloudflare/sandbox 0.12.10
@vercel/sandbox 3.5.0
@runloop/api-client 1.32.0
@declaw/sdk 1.5.0
@hopx-ai/sdk 0.5.1
@codesandbox/sdk 2.4.2
smolmachines 1.18.2
computesdk 4.1.8

Cloudflare's docs call the upcoming major API “1.0 preview,” while the registry's next tag resolves to 0.13.0-next.776.1. Keep the documentation label separate from the package version. CodeSandbox's pint tag is 2.5.0, but latest is 2.4.2; do not select a release merely by sorting version numbers. Source: Cloudflare docs; Cloudflare registry; CodeSandbox registry, checked 2026-09-25

Source: npm registry package metadata, checked 2026-09-25; exact package/tag snapshot preserved in the source review receipt.

Market map evidence

The local chart groups providers by statefulness and compute intensity. In the reviewed image, E2B, Vercel, and Cloudflare sit in the stateless/lightweight execution quadrant; Daytona, Runloop, declaw, CodeSandbox, and Hopx sit toward stateful lightweight environments; Modal, Northflank, and Tensorlake sit toward stateless compute-heavy workloads; Blaxel is nearer the center. Treat this as a market snapshot, not a taxonomy contract. It is useful because it makes each provider's implicit bet visible: short code execution, persistent devboxes, or heavier compute. Source: local image artifact, 2026-07-02

Isolation primitives and performance

Containers are lightweight and fast; microVMs give stronger isolation and are generally seen as the default "right" primitive; gVisor sits in between. Performance becomes multi-dimensional once a sandbox persists past one request: cold start dominates for code execution, resume time (snapshotting, warm pools) dominates for long-running sessions, and concurrency, disk, and workload density dominate at scale. Source: X/@crystalxtang, 2026-05-27

Local developer sandboxes become a first-class product

Docker Sandboxes turns the isolation pattern into a local developer workflow. The sbx CLI launches supported coding agents inside dedicated microVMs; each sandbox has its own Docker daemon, filesystem, and network while mounting only the project workspace into the environment. The agent can install packages, build containers, and modify its environment without receiving ambient access to the host. Docker also exposes editor integrations, an MCP gateway, reusable templates, and centrally managed network, filesystem, and MCP policies. The CLI is free for commercial use; organization governance is separately paid. Source: Hacker News item 49239751, 2026-08-10; Source: Docker Sandboxes documentation, 2026-08-10

This strengthens two existing decisions. First, local execution and strong isolation are no longer opposing product shapes: a desktop-native CLI can place the hardware boundary below the coding agent while preserving the familiar project-directory workflow. Second, agent portability increasingly depends on explicit workspace mounts, network policy, credential handling, and MCP policy, not merely on supporting the same agent binary. For Agent Machines and the environment-bootstrap workflow, Docker Sandboxes is therefore both a provider candidate and a contract reference; selection still requires startup, state, platform, policy, and cost comparison against persistent remote workers.

Apple's container adds a narrower macOS-native route. At commit 875d80ba07fdcc3aac138ed7fc56643d71aaf860, the Apache-2.0 project runs each Linux container as a lightweight VM, consumes and produces OCI images, is implemented in Swift, and exposes signed installers plus explicit upgrade, downgrade, keep-data, and delete-data paths. The current source requires Apple silicon and macOS 26, calls the project pre-1.0 with minor-version breakage possible, and contains 111 test files. Use it when a local Apple-silicon workflow wants OCI portability and VM isolation without a full remote sandbox; do not make it the universal agent worker because its host/platform boundary is deliberately narrow. Source: apple/container repository evidence at commit 875d80ba, captured 2026-08-12

Cloud Run Sandboxes and a locked-down VPS remain remote execution candidates, while Chromium alone is a browser process boundary rather than a general sandbox guarantee. Compare them on isolation, workspace persistence, egress, credentials, startup/resume, observability, cost, recovery, and the exact workload. A social claim that one route is "the sandbox we had the whole time" does not erase the control-plane boundary or prove containment. Source: X/@Vishal_anton16, 2026-07-12; X/@pk_iv, 2026-07-12; X/@zack_overflow, 2026-06-28

Filesystem as the agent memory layer

As agents accumulate context across files, logs, and cached data, the filesystem has become the default memory layer, so the execution environment stops being disposable and starts behaving like a persistent workspace. That demands snapshot and restore across sessions, branching and forking of execution state, and mounting of external storage (S3, blob, databases). This is the infra-side counterpart to Agent Memory Patterns. Source: X/@crystalxtang, 2026-05-27

Vercel Sandbox persistence is the mainstream platform version of that shift. A Vercel Sandbox is a long-lived named entity; each Session is one running VM. Stopping snapshots filesystem state, and resuming starts a new session from the latest snapshot. That design keeps the orchestration layer in control while preserving agent workspace state across compute sessions, which is why it belongs on the sandbox page even though it is not the same product shape as an always-on devbox. Source: Vercel changelog, 2026-03-26; Source: Vercel Sandbox docs, 2026-07-04

The 2026-06-16 duration change pushes that platform model further into long-running work. Vercel Sandboxes can now run uninterrupted sessions for up to 24 hours, up from 5 hours, with Sandbox.create({ timeout: 24 * 60 * 60 * 1000 }). Vercel frames the limit around large-scale data processing, E2E test pipelines, and long-lived agentic workflows, and explicitly recommends pairing the longer runtime with persistent sandboxes for durable state. The 24-hour maximum is available on Pro and Enterprise. Source: X/@vercel_dev, 2026-06-17; Source: Vercel changelog, 2026-06-16

Harness placement

The production default is now explicit: the trusted harness/control plane lives outside the replaceable execution sandbox. The sandbox is a tool-bearing worker reached through typed remote operations such as exec, read, write, browser control, and artifact transfer. It does not hold the authority or the state needed to recover from its own corruption. This is a trust and durability boundary, not merely a latency choice. Source: X/@NathanFlurry X Article, 2026-07-27

Trusted control plane, outside the sandbox Replaceable execution worker, inside the sandbox
Agent loop, retry ownership, cancellation, wake/resume, and run state Untrusted code, shells, build processes, browsers, dev servers, and transient caches
Canonical session/event history and cross-session index Checked-out workspace and recoverable execution-local files
Model credentials, provider routing, quotas, and cost policy Short-lived scoped capabilities; no ambient model or private-service secrets
Approvals, per-user authorization, tool policy, and agent-to-agent routing Tool requests whose authority is evaluated outside the worker
Audit log, observability, failure detection, schedules, and proof receipts Raw logs and artifacts emitted to the control plane with immutable references

If a runaway process exhausts memory, corrupts $PATH, or destroys the worker filesystem, recovery must still be able to observe the failure, retry from a checkpoint, and explain what happened. A persistent sandbox filesystem remains useful workspace memory, but it is not the canonical store for user history, approvals, policy, audit, or durable wiki knowledge. Persistence makes a worker convenient; it does not make the worker trusted.

Actor-like ownership is a strong implementation candidate for interactive, indefinitely lived agents: one addressable owner per agent can serialize work, hold local state, hibernate, wake on schedules, reconnect realtime clients, and restart after failure. Rivet's current Actors documentation exposes durable state, realtime, hibernation, queues, schedules, and SQLite, while its Flue beta maps each agent or workflow run to a durable actor and each execution context to an isolated agentOS VM with a persistent /workspace. That validates the split, not a universal mandate to adopt Rivet. Finite workflows may still fit a durable workflow engine; actor, database/queue, durable-object, or managed-agent implementations must be selected by replay cost, upgrade behavior, cancellation, realtime needs, portability, and measured recovery. The article's roughly 10ms same-datacenter tool-call overhead remains a vendor claim until reproduced on Kevin workloads. Source: Rivet Actors; Rivet Flue integration; agentOS Flue integration, reviewed 2026-08-10

Open Agents supplies a second, concrete implementation of this boundary without requiring an actor framework: a durable Workflow SDK agent loop persists chat, usage, cancellation, stream recovery, and post-finish steps outside a Vercel Sandbox that owns filesystem, shell, git, dev servers, snapshot, and hibernation. Its failure ledger adds operational constraints that belong in any provider adapter: snapshot is a stop transition; lifecycle timers must be durable; concurrent snapshots reconcile idempotently; observation endpoints stay read-only; and server lifecycle state wins over client clocks. This reinforces the contract while leaving the workflow/actor/queue choice open. Source: vercel-labs/open-agents at cf865e94, reviewed 2026-08-11

Execution substrate ladder

Do not ask only “which sandbox provider?” First decide how much operating-system surface the task actually needs:

Task contract Default execution class Examples Escalation boundary
Named service and data operations Typed capability or Code Mode executor query an API, transform records, call a validated host binding, bounded file reads/writes Shell, process, path, package, or POSIX behavior becomes part of correctness
Controlled coding-agent OS interop Agentos Virtual OS common shell tools, git from the registry, Node/Python, virtual processes, mounted storage, static build artifacts
General or hostile computer workload Full sandbox, microVM, persistent devbox, or desktop Playwright/browser, native toolchains, databases and hot reload, Docker, desktop automation, unreviewed binaries Work can return to a narrower tier with a content-addressed checkpoint

The virtual-OS insight is that shell, filesystem, process, and network semantics can be implemented without provisioning a conventional Linux machine for every operation. The stronger durable rule is capability escalation: start at the narrowest class that preserves the program's semantics and threat model, then escalate only the unsupported operation. One logical workspace can be mounted or overlaid across tiers, but the receipt records the implementation/version, permissions, mounts, secret path, limits, escalation reason, state transfer, provider, cost, and proof. Source: X/@NathanFlurry, “you probably don't need an expensive sandbox,” 2026-06-29; Source: agentOS versus Sandbox and Sandbox Mounting, reviewed 2026-08-11

Security does not order these classes globally. A virtual kernel can offer finer programmatic filesystem, process, and egress policy than a coarse box, while a separately hardened microVM can offer a smaller blast radius than an in-process runtime. agentOS currently calls its security model beta, assigns host hardening and policy to the operator, and excludes arbitrary native software and several kernel/hardware surfaces. Therefore the source's blanket “more secure than a sandbox” claim is held. Choose by adversary, trusted computing base, shared failure domain, required capabilities, and reproduced escape/resource tests—not by category name. Source: agentOS Security Model and Limitations, reviewed 2026-08-11

Two implementation cautions follow. First, a virtual command must not run agent-generated build code directly in the trusted host merely because the output is static; execute it inside the selected isolation tier or a narrowly reviewed binding. Second, snapshot and fork are often filesystem operations, but workspace branching does not branch canonical user history, approvals, policy, or proof. Those remain control-plane records with explicit parentage.

Collaborative rooms multiply sandbox authority

A multiplayer coding room is not merely a sandbox with more cursors. Every participant can become an operator of the same terminal, files, agent, provider budget, and preview surface; the room link, repository connection, credentials, and public preview are separate authority edges. The control plane must record who can observe, type, approve, connect a repository, introduce a secret, spend, export, invite, revoke, delete, and publish. Joining collaboration must never silently broaden a worker's repository, secret, or network scope.

Jam is a useful current reference and guarded candidate for this shape. Its published contract uses email magic-code auth, BYOK Anthropic access sent from an E2B sandbox, InstantDB collaboration state, shared terminal visibility, ephemeral sandbox retention, and optional persistent public project/preview URLs. The replay corroborated those client-visible surfaces but did not create an account or prove server-side isolation, authorization, deletion, or provider configuration. Route Jam only for a disposable fork or an intentionally shared workspace with least-privilege BYOK credentials, an explicit participant/repository manifest, no ambient local secrets, bounded spend/egress, revocation and deletion checks, and a separate approval for public preview. Keep private or production-bearing work on the existing governed local/remote execution path until those controls are proved. Source: X 2070242834985431293; Jam live UI, privacy policy, terms, public client schema, and browser replay, 2026-08-11

The same boundary is starting to appear in CI/CD product ideas. Brendan Irvine-Broque's "CI/CD from first principles" bookmark describes CI as a workflow plus a sandbox: TypeScript orchestration instead of opaque GitHub Actions YAML, local debugging, durable execution, retries, and VM/directory snapshots. There is no source artifact beyond the tweet, so this is not a tool endorsement; it is a concise statement of the product direction where CI becomes a debuggable workflow runtime over persistent execution state. Source: X/@irvinebroque, 2026-06-21

Provider bets

Provider comparison uses one manifest and one hostile-but-authorized fixture, not incomparable launch numbers. Freeze host and guest OS/architecture, hypervisor/runtime and image digest, tenant identity, mounts and persistence, secret injection, ingress/egress, package mirrors, privileges, resource limits, snapshot/fork/resume semantics, patching, logging, deletion, and recovery. Then measure cold start, warm resume, steady runtime, burst creation, failure and retry behavior, cost, escape and exfiltration canaries, and operator recovery. Modal's million-sandbox article, Fly Sprites, AWS AgentCore, smolvm, and HN microVM discussions are useful hypotheses only when run through this same contract. Published fleet scale or a microVM label does not establish tenant isolation, a compatible workload, or recoverability. Source: X Articles 2077844613428031488, 2080672317123047424, 2075676517674418176, and 2080484722204131328; smol-machines/smolvm@f72f9401; Hacker News items 49240545, 49243972, and 49245030, reviewed 2026-08-12

Each major provider optimizes for a different agentic future: Modal treats sandboxes as one workload on a broader compute layer (betting on traditional infra shapes), e2b uses Firecracker microVMs and bets most execution stays stateless, and Daytona was described as a container-based devbox betting on persistent operating environments. The current Daytona product distinguishes container, VM, and GPU classes; use Daytona for that contract. At enough scale, companies stop buying and build in-house: the baseline is commoditizing, cold starts are converging, and the real competition is hyperscalers (AWS already exposes Lambda, ECS, and Firecracker). Source: X/@crystalxtang, 2026-05-27

Rivet agentOS is the WebAssembly/V8 virtual-kernel counter-position: instead of provisioning a full sandbox per task, run controlled filesystem, process, network, and shell semantics in a compact runtime and escalate through sandbox mounting when the declared surface is insufficient. Its performance and cost figures remain workload-specific vendor benchmarks; its current security model is beta. Use named bindings first for bounded operations, agentOS for OS interop within its registry and limits, and full workers for arbitrary native, browser/GUI, kernel, hardware, or higher-risk work. Source: X/@rivet_dev, 2026-07-04; Source: current agentOS repository and documentation, reviewed 2026-08-11

A newer shape worth tracking: Orgo sells computer-use desktops — full GUI VDI machines agents drive by screenshot, sub-500ms boot — rather than code sandboxes, a distinct bet from the Modal/e2b/Daytona code-execution lineage. Source: orgo.ai, 2026-06-15

Above the providers sits a unification layer: ComputeSDK gives one typed API over 14+ of them (E2B, Modal, Daytona, Vercel, Cloudflare, …), so app code is provider-agnostic — the sandbox analog of the Vercel AI SDK / OpenRouter for models, and the open-source shape of Agent Machines' substrate routing. Source: computesdk.com, 2026-06-15

CubeSandbox is the notable self-hosted microVM candidate from the current X replay. Its source describes a RustVMM/KVM runtime with an E2B-compatible API, sub-60ms cold-start and under-5MB overhead claims, a per-sandbox eBPF network boundary, an L7 egress proxy that can inject credentials without exposing them to guest code, snapshot/clone/rollback, AutoPause/AutoResume, volumes, and single-node, Terraform, or Kubernetes deployment paths. That combination makes it relevant to both Agent Machines and the workflow sandbox contract: secrets remain outside workers, egress is policy, and state can be checkpointed or branched. It does not become the default from a README benchmark. Adoption must reproduce latency and density on the target hardware, resolve the quick-start/platform matrix, measure E2B compatibility gaps, review the still-preview cluster features, and prove credential, egress, snapshot, upgrade, recovery, and audit behavior under hostile workloads. Source: X 2076504402308047281; TencentCloud/CubeSandbox README and repository snapshot at b2e1fad, reviewed 2026-08-10

AgentENV is retained as a conditional self-hosted Firecracker/KVM platform when an E2B-compatible API and large Linux microVM fleet are the named job. The pinned source is MIT-licensed, has tests and CI, targets Ubuntu 24.04/Linux 6.8+ with /dev/kvm, and claims sub-100ms lifecycle operations. It is not a Mac install and the upstream warning that authorization is not yet implemented is disqualifying for public exposure. Evaluate only on a dedicated Linux/KVM host behind loopback, a trusted private network, or an independently authenticated/mTLS proxy. Before promotion, prove tenant isolation, least privilege, image/OCI provenance, egress and secret injection, quotas, lifecycle/delete, encrypted snapshot storage, backup/restore, upgrade/rollback, audit, measured boot/resume/pause/snapshot distributions, and E2B compatibility. Never pipe its privileged installer into a host without reviewing the pinned scripts and rollback. Source: X 2081762978391843020; kvcache-ai/AgentEnv@6e8deeab, reviewed 2026-08-12

Box by ASCII is retained as a guarded managed full-VM candidate for work that genuinely needs persistent Ubuntu, SSH/SCP, Docker inside the worker, an interactive desktop, dedicated IPv4, public HTTPS, or snapshot/fork/template lifecycle. Its current default machine is 4 shared vCPUs and 8 GB at a documented $0.036/hour; stopped boxes do not consume machine time, while create, fork, and resume each consume a start. Those properties make it interesting for long-running agent platforms, but not a universal replacement for a typed capability, virtual OS, or an already integrated provider. Regions are currently EU-only, and the published ceilings of 600 starts/hour and 1,500 starts/day must be part of capacity design. The X video is a documentation tour, not proof of desktop frame rate, fork latency, isolation, or production recovery. Source: Box platform guide, billing, FAQ, and X 2076651776527224913, reviewed 2026-08-11

Admit Box through this provider-specific boundary:

  1. Use the official HTTP API or an exact pinned SDK for the first controlled pilot. The reviewed shell installer downloads an opaque CLI binary, changes shell startup files, and starts interactive onboarding without verifying a published checksum or signature, so do not pipe it into a shell or install it globally by default.
  2. Create untrusted-user workers from a clean environment already marked safeForThirdParties, or send noEnv: true on creation. Do not treat conversion of a previously trusted snapshot as complete scrubbing: Box removes credentials it manages, but explicitly leaves operator-added AWS, gcloud, .netrc, .npmrc, and Docker credentials in place. Build reusable templates without secrets; pass only task-scoped capabilities afterward.
  3. Keep the account bearer key in the trusted control plane. The current v1 schema documents account-scoped bearer authentication and no resource scopes; workers receive neither that key nor ambient provider credentials. Bind every box to tenant, purpose, TTL, spend, start/concurrency budget, source template, and teardown receipt.
  4. Treat desktop and hosted URLs as separate authority surfaces. Never put a protected URL token in logs, screenshots, or referrers, and require a distinct approval before making desktop or service access public. Dedicated IPv4 and broad TCP/UDP reach increase the need for explicit ingress and egress policy; they do not prove a particular hypervisor or containment strength.
  5. Keep sensitive, regulated, production, and third-party platform use blocked pending policy resolution. The Terms require written approval before reselling or providing third-party access to Box compute, disclaim production/high-risk suitability, permit operational monitoring and retention, and link to ariana.dev/privacy, which currently redirects to a 404 at ascii.dev/privacy. Obtain the applicable privacy terms and written platform-use approval before that boundary is crossed.

Source: environments and no-env behavior, v1 API, hosting, and Terms, captured 2026-08-11

Relevance to Kevin

Grounds the agent-infra layer behind Open Agents.dev (Sandboxes plus Workflow SDK plus AI Gateway) and the Vercel Sandbox skill. AI SDK 7 HarnessAgent pairs harness adapters with createVercelSandbox() — see AI SDK HarnessAgent. The "sandbox as agent OS" framing reinforces Agent Operating System: persistent, stateful execution is becoming the substrate, not a disposable utility.


Timeline