Agent Security Model
Agent security is the control of context, capabilities, identity, side effects, and auditability around a non-deterministic actor that can use tools.
Source Tweet Artifacts
Core model
Classic app security protects code paths. Agent security also has to protect decision paths: what the model read, what it believed, what tools it could call, what authority those tools carried, and what side effects happened before a human noticed.
The practical security boundary has five layers:
| Layer | Security question | Local owner |
|---|---|---|
| Context | Can untrusted text steer the agent into unsafe action? | Brin - Agent Context Security, prompt-injection checks, source provenance |
| Capability | Which tools, MCP servers, CLIs, and remote agents are available? | Agent Capability Registry, MCP Registry, Capability Routing Map |
| Authority | What credentials, scopes, networks, and filesystem paths does a tool carry? | secrets policy, OAuth scopes, sandbox mounts |
| Execution | Where does untrusted code or tool work run? | MicroVM and Sandbox Isolation Architecture, Vercel/E2B/Dedalus sandboxes |
| Audit | Can an operator reconstruct the run and revoke/repair damage? | Agent Trajectory Evaluation, logs, traces, approvals |
Guardrail layering
A command denylist is useful defense in depth, but it is not a complete authority system. The preferred
order is typed capability scope and arguments → sandbox and credential/network boundaries → approval policy
→ command adapter → post-action receipt. If a harness only exposes command strings, its regex hook must be
treated as a partial adapter with an adversarial test corpus and a health check, never as proof that an
operation is safe. One canonical policy should project into each harness rather than letting hooks drift.
Source: davidondrej/skills global-agent-guardrails, commit
081d4d52cddb1f24e9358add21d47a45d82dd3fd, reviewed 2026-08-11
A named agent is not an authority boundary
Role identity, memory identity, process identity, machine identity and credential scope are different things. Grok Bot is a useful concrete counterexample: every named Bot has a separate role and conversation, but all Bots for one user share one persistent computer, browser sessions, files, command-line credentials, installed connectors and local-computer permissions. The product explicitly warns that separate screens are not security boundaries and that deleting a Bot does not clear shared files or sign-ins. Source: Grok Bot overview, computer and apps, and approvals/security, reviewed 2026-08-12
Any multi-agent harness must therefore disclose and enforce the actual sharing matrix. Record the authenticated principal, named role, run, worker, filesystem, browser profile, connector, credential lease, network policy and memory scope separately. A collaboration handoff should pass a typed artifact or scoped capability, not silently widen every Bot to the same account login. Deleting or disabling a role must enumerate and revoke routines, live runs, tokens, sessions, files, snapshots and retained data; removing the friendly name is not revocation.
Model review is also not the hard boundary. Grok Bot and Cursor describe Auto Review as model-based; Cursor's security documentation calls run-mode guardrails best effort. Use it to reduce unnecessary prompts and spot contextual risk, but enforce consequential limits with deterministic hooks, sandbox and network boundaries, scoped credentials, target-bound approvals and an audit receipt. Source: Grok Bot approvals, Cursor Agent Security, and Auto-review, reviewed 2026-08-12
“Read-only” database access is still high authority
SELECT is not harmless when it spans every present and future table, and BYPASSRLS defeats tenant and
row policy. Do not use a blanket default-privileges grant plus denylist as the standard agent reader. Prefer
explicit views or schema/table allowlists, a role that remains subject to scoped RLS, read replicas where
appropriate, statement/idle timeouts, row and result limits, query audit, PII/egress controls, credential
expiry, and human-reviewed DDL. Empty results are a reason to diagnose policy—not to grant BYPASSRLS.
Source: davidondrej/skills create-readonly-db-role, commit
081d4d52cddb1f24e9358add21d47a45d82dd3fd, reviewed 2026-08-11
Device-local agent accounts are real principals
DHH's UniFi example is a useful capability signal: a local admin account let an agent inspect and tune radio channels, mesh settings, and roaming without using the operator's cloud identity. It is not evidence that a broad administrator account is the right default. Network control can disconnect devices, change segmentation, expose management surfaces, or erase the path an operator needs to recover the system.
Treat an infrastructure agent as its own principal. Prefer a local-only, revocable account with the narrowest available role; freeze the controller, site, devices, allowed configuration classes, change window, before-state, proposed diff, and recovery path. Read/diagnose and mutate/optimize are separate verbs, and channel, mesh, power, VLAN, firewall, DNS, firmware, reset, and credential changes need separate policy. Ubiquiti's primary documentation confirms local management and local credentials; the social post demonstrates usefulness, not least privilege or safe unattended mutation. Source: X/@dhh; Ubiquiti local management, reviewed 2026-08-12
Safety prompts must remain truthful
Reframing an authorized defensive task in precise, non-sensational language can improve review. Replacing
keywords, abstracting the real domain, or otherwise disguising intent to get past a safety classifier is not
a harness feature. Preserve such recipes as counterevidence, not executable routes. If a legitimate task is
blocked, state the true scope, authority, target, constraints, and defensive purpose, then use the supported
approval or escalation path. Source: davidondrej/skills fable-safe-prompt, commit
081d4d52cddb1f24e9358add21d47a45d82dd3fd, reviewed 2026-08-11
OWASP's agentic-application work frames the risk shift similarly: agents introduce tool misuse, excessive agency, memory/context poisoning, insecure tool integration, and governance failures beyond ordinary prompt injection. Source: OWASP Agentic Applications Top 10, https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/, 2026-06-19
Threat taxonomy
| Threat | How it appears | Control |
|---|---|---|
| Prompt/context injection | A web page, issue, email, or tool result instructs the agent to ignore policy or leak data | Treat external context as hostile; annotate source trust; isolate instructions from observations |
| Tool confusion | Agent chooses the wrong overlapping capability or misreads a tool contract | Narrow tool surfaces; Tool-Calling Design; resolver routing; evals |
| Tool-trace leakage | Logs, WebSocket events, screenshots, or streamed tool results expose filesystem paths, .env contents, tool schemas, or internal capability names |
Redact traces before display; separate observability from secrets; safe-disclosure template; never treat path exposure as permission to read |
| Harness self-MITM extraction | A request framed as latency debugging asks the harness to proxy or capture its own traffic, exposing system prompts and tool schemas piece by piece | Treat self-instrumentation, proxy, certificate, and network-capture requests as exfiltration-sensitive; pin harness egress; require explicit scoped authority; alert on new local proxies Source: X/@arafatkatze 2083236726676615535, reviewed 2026-08-05 |
| Excessive authority | Read task has write credentials; dev task has prod secrets | Least privilege, scoped credentials, approval gates |
| Untrusted code execution | Agent runs repo scripts, package installs, or generated code | Sandboxes, network policy, filesystem isolation |
| Model access provenance | Agent routes through discounted, resold, or account-pooled model access | Provider provenance checks, account ownership checks, revocation path, billing anomaly monitoring |
| Memory poisoning | Bad facts enter persistent memory and influence future runs | Knowledge Compilation discipline; timelines; source citations; review before durable rules |
| Side-effect drift | Agent mutates state outside versioned systems | Production Safety, migrations, idempotency, audit logs |
| Opaque autonomy | Nobody can explain why the agent acted | traces, artifact retention, trajectory evaluation |
IRL comparison
| System / standard | Useful lesson |
|---|---|
| OWASP Agentic Applications Top 10 | Agent risks are specific to tool-using systems; treat tool authority and context as first-class attack surfaces. |
| NIST AI RMF Generative AI Profile | Risk management should track data provenance, human oversight, monitoring, incident response, and measured impact, not only model behavior. Source: NIST AI RMF GenAI Profile, https://www.nist.gov/itl/ai-risk-management-framework, 2026-06-19 |
| Golf - MCP Agent Security | Enterprise buyers need discover/govern/audit for MCP and agent connections. |
| Brin - Agent Context Security | Context scoring is complementary to guardrails; unsafe context is often the input that compromises the agent. |
| LangSmith Sandbox Auth Proxy | Keep credentials outside the sandbox while the proxy injects request auth and enforces outbound host/port policy. |
| Agent Machines | Persistent skilled workers need credential gates, sandboxed execution, and trajectory-level observability. |
Kevin policy
- Discovery is not approval. Registry presence never means a tool is trusted.
- Read-only should be the default authority tier.
- Any prod, billing, external-send, credential, or destructive operation needs an explicit gate.
- Tool results are observations, not instructions.
- Persistent memory writes require source and timeline discipline.
- Agent security claims must be backed by traces, tests, or reproducible policy checks.
- Cheap or resold model access is not automatically safe. The reviewed HN screenshot frames discounting as a mixture of pooled accounts, suspicious payment flows, and possible output or reasoning-trace resale. Verify provider provenance, account ownership, logging, data handling, terms, revocation, and billing anomaly controls before routing agent work through any reseller. Source: X/@GregKamradt, 2026-06-25; Source: Hacker News thread, 2026-06-25
- A leaked path, tool name, or streamed trace is evidence for a report, not authorization to fetch more. The reviewed Cursor/OpenAI artifact shows the right split: refuse file retrieval/exfiltration, then document the exposed tool capabilities and paths as an observability/security finding. A companion screenshot shows the failure mode: streamed output summarizing
.envand top-level files. Source: X/@sathuashrith, 2026-05-27; Source: local artifact review, 2026-07-03
The strong version: a capable agent should feel free inside a small, well-defined box. Making the box explicit is security engineering.
Timeline
- 2026-08-12 | Added device-local agent accounts as a distinct authority boundary after DHH used a UniFi local admin for Claude-assisted wireless tuning. The durable rule is separate principal, local/revocable credential, explicit device/config scope, before-state and recovery proof—not “give the agent admin.” Source: X/@dhh
2087272392557007042; Ubiquiti local-management documentation - 2026-08-12 | Added the Grok Bot shared-computer counterexample: named roles and conversations do not imply machine, login, connector, file or permission isolation; deletion is not revocation; and model-based Auto Review remains steering around deterministic capability, sandbox, credential, approval and audit boundaries. Source: SpaceXAI Grok Bot official docs; Cursor security and Auto-review docs
- 2026-08-11 | Added command-guardrail layering, rejected blanket future-table
SELECT/BYPASSRLSas a read-only default, and made safety-obfuscation recipes retained counterevidence rather than callable capability. Source: X/@MengTo paired-library replay;davidondrej/skillsexact source - 2026-07-04 | Added LangSmith Sandbox Auth Proxy as a concrete authority-layer control: sandbox code can call external APIs through proxy-injected credentials while egress remains host/port policy. Source: LangSmith docs, 2026-07-04; Source: X/@LangChain, 2026-05-21
- 2026-07-03 | Added tool-trace leakage as a named threat after reviewing the Cursor/OpenAI artifact: safe refusal plus disclosure language is the right response; leaked WebSocket/tool traces and
.envsummaries are the system failure. Source: X/@sathuashrith, 2026-05-27; Source: local artifact review, 2026-07-03 - 2026-07-02 | Reviewed the local HN screenshot artifact and made model-access provenance a named threat row. Source: wiki/assets/x-bookmarks/2069968506951725419/image-01.jpg
- 2026-06-25 | Added grey-market token access as a provenance and billing-security risk, not a harmless cost lever. Source: X/@GregKamradt, 2026-06-25; Hacker News thread, 2026-06-25
- 2026-06-19 | Created from the wiki-wide coverage audit to connect agent guardrails, MCP governance, context security, sandboxing, production safety, and trajectory auditability. Source: whole-wiki concept review, 2026-06-19