Agent Security Model

Agent security is the control of context, capabilities, identity, side effects, and auditability around a non-deterministic actor that can use tools.

Source Tweet Artifacts

Core model

Classic app security protects code paths. Agent security also has to protect decision paths: what the model read, what it believed, what tools it could call, what authority those tools carried, and what side effects happened before a human noticed.

The practical security boundary has five layers:

Layer Security question Local owner
Context Can untrusted text steer the agent into unsafe action? Brin - Agent Context Security, prompt-injection checks, source provenance
Capability Which tools, MCP servers, CLIs, and remote agents are available? Agent Capability Registry, MCP Registry, Capability Routing Map
Authority What credentials, scopes, networks, and filesystem paths does a tool carry? secrets policy, OAuth scopes, sandbox mounts
Execution Where does untrusted code or tool work run? MicroVM and Sandbox Isolation Architecture, Vercel/E2B/Dedalus sandboxes
Audit Can an operator reconstruct the run and revoke/repair damage? Agent Trajectory Evaluation, logs, traces, approvals

Guardrail layering

A command denylist is useful defense in depth, but it is not a complete authority system. The preferred order is typed capability scope and arguments → sandbox and credential/network boundaries → approval policy → command adapter → post-action receipt. If a harness only exposes command strings, its regex hook must be treated as a partial adapter with an adversarial test corpus and a health check, never as proof that an operation is safe. One canonical policy should project into each harness rather than letting hooks drift. Source: davidondrej/skills global-agent-guardrails, commit 081d4d52cddb1f24e9358add21d47a45d82dd3fd, reviewed 2026-08-11

A named agent is not an authority boundary

Role identity, memory identity, process identity, machine identity and credential scope are different things. Grok Bot is a useful concrete counterexample: every named Bot has a separate role and conversation, but all Bots for one user share one persistent computer, browser sessions, files, command-line credentials, installed connectors and local-computer permissions. The product explicitly warns that separate screens are not security boundaries and that deleting a Bot does not clear shared files or sign-ins. Source: Grok Bot overview, computer and apps, and approvals/security, reviewed 2026-08-12

Any multi-agent harness must therefore disclose and enforce the actual sharing matrix. Record the authenticated principal, named role, run, worker, filesystem, browser profile, connector, credential lease, network policy and memory scope separately. A collaboration handoff should pass a typed artifact or scoped capability, not silently widen every Bot to the same account login. Deleting or disabling a role must enumerate and revoke routines, live runs, tokens, sessions, files, snapshots and retained data; removing the friendly name is not revocation.

Model review is also not the hard boundary. Grok Bot and Cursor describe Auto Review as model-based; Cursor's security documentation calls run-mode guardrails best effort. Use it to reduce unnecessary prompts and spot contextual risk, but enforce consequential limits with deterministic hooks, sandbox and network boundaries, scoped credentials, target-bound approvals and an audit receipt. Source: Grok Bot approvals, Cursor Agent Security, and Auto-review, reviewed 2026-08-12

“Read-only” database access is still high authority

SELECT is not harmless when it spans every present and future table, and BYPASSRLS defeats tenant and row policy. Do not use a blanket default-privileges grant plus denylist as the standard agent reader. Prefer explicit views or schema/table allowlists, a role that remains subject to scoped RLS, read replicas where appropriate, statement/idle timeouts, row and result limits, query audit, PII/egress controls, credential expiry, and human-reviewed DDL. Empty results are a reason to diagnose policy—not to grant BYPASSRLS. Source: davidondrej/skills create-readonly-db-role, commit 081d4d52cddb1f24e9358add21d47a45d82dd3fd, reviewed 2026-08-11

Device-local agent accounts are real principals

DHH's UniFi example is a useful capability signal: a local admin account let an agent inspect and tune radio channels, mesh settings, and roaming without using the operator's cloud identity. It is not evidence that a broad administrator account is the right default. Network control can disconnect devices, change segmentation, expose management surfaces, or erase the path an operator needs to recover the system.

Treat an infrastructure agent as its own principal. Prefer a local-only, revocable account with the narrowest available role; freeze the controller, site, devices, allowed configuration classes, change window, before-state, proposed diff, and recovery path. Read/diagnose and mutate/optimize are separate verbs, and channel, mesh, power, VLAN, firewall, DNS, firmware, reset, and credential changes need separate policy. Ubiquiti's primary documentation confirms local management and local credentials; the social post demonstrates usefulness, not least privilege or safe unattended mutation. Source: X/@dhh; Ubiquiti local management, reviewed 2026-08-12

Safety prompts must remain truthful

Reframing an authorized defensive task in precise, non-sensational language can improve review. Replacing keywords, abstracting the real domain, or otherwise disguising intent to get past a safety classifier is not a harness feature. Preserve such recipes as counterevidence, not executable routes. If a legitimate task is blocked, state the true scope, authority, target, constraints, and defensive purpose, then use the supported approval or escalation path. Source: davidondrej/skills fable-safe-prompt, commit 081d4d52cddb1f24e9358add21d47a45d82dd3fd, reviewed 2026-08-11

OWASP's agentic-application work frames the risk shift similarly: agents introduce tool misuse, excessive agency, memory/context poisoning, insecure tool integration, and governance failures beyond ordinary prompt injection. Source: OWASP Agentic Applications Top 10, https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/, 2026-06-19

Threat taxonomy

Threat How it appears Control
Prompt/context injection A web page, issue, email, or tool result instructs the agent to ignore policy or leak data Treat external context as hostile; annotate source trust; isolate instructions from observations
Tool confusion Agent chooses the wrong overlapping capability or misreads a tool contract Narrow tool surfaces; Tool-Calling Design; resolver routing; evals
Tool-trace leakage Logs, WebSocket events, screenshots, or streamed tool results expose filesystem paths, .env contents, tool schemas, or internal capability names Redact traces before display; separate observability from secrets; safe-disclosure template; never treat path exposure as permission to read
Harness self-MITM extraction A request framed as latency debugging asks the harness to proxy or capture its own traffic, exposing system prompts and tool schemas piece by piece Treat self-instrumentation, proxy, certificate, and network-capture requests as exfiltration-sensitive; pin harness egress; require explicit scoped authority; alert on new local proxies Source: X/@arafatkatze 2083236726676615535, reviewed 2026-08-05
Excessive authority Read task has write credentials; dev task has prod secrets Least privilege, scoped credentials, approval gates
Untrusted code execution Agent runs repo scripts, package installs, or generated code Sandboxes, network policy, filesystem isolation
Model access provenance Agent routes through discounted, resold, or account-pooled model access Provider provenance checks, account ownership checks, revocation path, billing anomaly monitoring
Memory poisoning Bad facts enter persistent memory and influence future runs Knowledge Compilation discipline; timelines; source citations; review before durable rules
Side-effect drift Agent mutates state outside versioned systems Production Safety, migrations, idempotency, audit logs
Opaque autonomy Nobody can explain why the agent acted traces, artifact retention, trajectory evaluation

IRL comparison

System / standard Useful lesson
OWASP Agentic Applications Top 10 Agent risks are specific to tool-using systems; treat tool authority and context as first-class attack surfaces.
NIST AI RMF Generative AI Profile Risk management should track data provenance, human oversight, monitoring, incident response, and measured impact, not only model behavior. Source: NIST AI RMF GenAI Profile, https://www.nist.gov/itl/ai-risk-management-framework, 2026-06-19
Golf - MCP Agent Security Enterprise buyers need discover/govern/audit for MCP and agent connections.
Brin - Agent Context Security Context scoring is complementary to guardrails; unsafe context is often the input that compromises the agent.
LangSmith Sandbox Auth Proxy Keep credentials outside the sandbox while the proxy injects request auth and enforces outbound host/port policy.
Agent Machines Persistent skilled workers need credential gates, sandboxed execution, and trajectory-level observability.

Kevin policy

  • Discovery is not approval. Registry presence never means a tool is trusted.
  • Read-only should be the default authority tier.
  • Any prod, billing, external-send, credential, or destructive operation needs an explicit gate.
  • Tool results are observations, not instructions.
  • Persistent memory writes require source and timeline discipline.
  • Agent security claims must be backed by traces, tests, or reproducible policy checks.
  • Cheap or resold model access is not automatically safe. The reviewed HN screenshot frames discounting as a mixture of pooled accounts, suspicious payment flows, and possible output or reasoning-trace resale. Verify provider provenance, account ownership, logging, data handling, terms, revocation, and billing anomaly controls before routing agent work through any reseller. Source: X/@GregKamradt, 2026-06-25; Source: Hacker News thread, 2026-06-25
  • A leaked path, tool name, or streamed trace is evidence for a report, not authorization to fetch more. The reviewed Cursor/OpenAI artifact shows the right split: refuse file retrieval/exfiltration, then document the exposed tool capabilities and paths as an observability/security finding. A companion screenshot shows the failure mode: streamed output summarizing .env and top-level files. Source: X/@sathuashrith, 2026-05-27; Source: local artifact review, 2026-07-03

The strong version: a capable agent should feel free inside a small, well-defined box. Making the box explicit is security engineering.


Timeline