Security Scan Workflow
On-demand Security and Review Skills security scans investigate a designated repo with explicit cost approval, verified findings, and post-scan documentation.
Kevin must invoke explicitly: "Run the security-scan automation against <repo>." This workflow is manual-only because full deepsec processing can be expensive and high-impact. Source: automations/security-scan.md, 2026-05-31
Routing
Use this workflow for a full repo security scan. Use ordinary code review for small diffs, Code Bugfix Workflow for recent author-owned regressions, and postmortem pages only after a verified issue or material lesson exists. For an Express service, treat Helmet as a cheap application-hardening and verification layer before the expensive repository-wide scan—not as a replacement for either code review or Deepsec.
Express/Helmet preflight
At pinned Helmet commit 9315aac, app.use(helmet()) sets 13 response headers,
including a default Content Security Policy and HSTS. That default is a starting
point, not a proof: Helmet performs little CSP validation, its CSP often needs
application-specific directives/nonces, upgrade-insecure-requests can break
Safari localhost development, and Cross-Origin-Embedder-Policy is not enabled
by default. For an Express target:
- freeze the deployed origin, routes, assets, embeds, workers, APIs, auth flows, and development versus production behavior;
- configure Helmet explicitly, start CSP changes in Report-Only where needed, and avoid weakening directives just to silence failures;
- assert actual response headers on representative routes and states, then run a CSP evaluator and browser smoke tests for scripts, styles, fonts, images, forms, frames, downloads, OAuth, error pages, and localhost behavior;
- keep the resulting header receipt with the normal dependency, secret, authorization, data-flow, and repository scan evidence.
Helmet is MIT, mature, and well tested (40 test files in the pinned tree), but headers cannot repair application authorization, injection sinks, leaked secrets, vulnerable dependencies, or unsafe business logic.
Authorized wireless lab branch
Repository scanning does not authorize live network testing. When the target is
a wireless network, freeze a separate signed scope containing owner, SSIDs,
BSSIDs, physical location, client/device exclusions, time window, operator,
hardware, evidence retention, notification, restore, and emergency-stop
contact. Begin with passive inventory. Airgorah may be used on a dedicated Linux
lab host and monitor-mode adapter only after that scope exists; it requires root
and can capture traffic, deauthenticate clients, capture handshakes, and attempt
password cracking. Each disruptive action needs an explicit second approval and
must stop on an out-of-scope radio/client, service impact, ambiguous ownership,
or missing restore path. Never run this branch against a network merely because
it is reachable. Source: martin-olivier/airgorah@2acaa3b, 2026-08-12
Intentionally vulnerable AI training range
LLMVault is retained as a local, authorized training fixture for prompt
injection, indirect/multimodal injection, RAG authorization failures, unsafe
output handling, excessive agency, leakage, poisoning, extraction, and
unbounded-consumption exercises. It is deliberately insecure, and its own
documentation says not to expose it to the public internet or reuse the code in
production. Run a pinned commit on loopback inside a disposable container/VM,
with no real credentials, private corpora, production models, unrestricted
egress, or routable bridge; prefer scripted Play Mode before optional local
Ollama Live Mode. Snapshot/reset its progress data, preserve attack and defense
receipts, and destroy the environment after the lab. The MIT code license does
not license the LLMVault branding/artwork. Three test files and one CI workflow
support training use, not isolation assurance. Source: CyberSunil/LLMVault@0fa0ce7, 2026-08-12
Model-surgery and refusal-removal research lane
Tools such as Obliteratus that locate or alter refusal-related model weights are
retained as research evidence, not production routing. Any run needs an
authorized model and dataset, exact weights/tokenizer/code digests, license and
provenance, an isolated offline worker, no operational credentials or external
targets, reversible checkpoints, before/after capability and safety evals,
misuse and dual-use review, and a separately approved publication/export plan.
The output is a bounded research receipt; it never silently replaces a served
model or weakens the ordinary safety policy. A social demo proves neither that
the edited mechanism is understood nor that general capabilities and safeguards
remain intact. Source: X 2080607945071686030, reviewed 2026-08-12
Framework advisory intake
A saved fixed-version announcement is a trigger, never a permanent minimum.
Resolve the target's installed version and feature exposure, fetch the current
maintainer advisories, compute the highest applicable fixed line, update the
lockfile, inspect transitive duplicates, and exercise affected routes plus
auth/cache/proxy/server-action/image/WebSocket behavior. For the May 2026
Next.js announcement, 16.2.6 / 15.5.18 fixed that disclosed set; July
advisories subsequently raised the applicable maintained-line floors to
16.2.11 / 15.5.21. Current authority must always win over the social post.
Source: vercel/next.js security advisories and releases, checked 2026-08-12
Agent boundary and credential preflight
Agent security has at least four independent boundaries. Passing one does not prove the others:
| Boundary | Required proof |
|---|---|
| Skill/package supply chain | Pinned source tree, license, dependencies, scripts, hooks, network/global writes, static scan, semantic-intent review, tests, signature/provenance where available, and independent manual review. NVIDIA SkillSpector is a useful Apache-2.0 scanner with static and optional semantic checks; it is one layer, not an install verdict. |
| Credential mediation | The sandbox receives a narrow capability or proxy route, not the raw provider secret; prove provider/tenant/tool identity, scopes, audience, expiry, revocation, request logging, response filtering, rate/cost limits, and confused-deputy resistance. LangSmith's Auth Proxy, DAuth, and Treg are patterns/candidates, not interchangeable proof. |
| Runtime containment | Freeze filesystem, network/egress, package mirrors, secrets, subprocesses, persistence, lifecycle, audit, and escape test. A command denylist or approval prompt is not a hostile-code sandbox. |
| External authority | Reads, drafts, deploys, purchases, sends, deletes, and publishes are different verbs. Bind each consequential action to exact target, account, parameters, content/proposal digest, actor, approval, idempotency key, result, and rollback/recovery receipt. |
File identification and document conversion are preflights, not containment.
Magika can supply a content-type hypothesis before parser selection, and
MarkItDown can supply a text/structure projection, but neither proves that a
polyglot, truncated, mislabeled, malicious, or parser-triggering file is safe.
Preserve the original digest; compare extension, claimed type, detector output,
and parser result; route ambiguity or disagreement into an isolated parser;
retain page/slide/sheet/time locators, omitted media and structures, failures,
network/provider use, and deletion state. Malware scanning, decompression and
resource limits, parser sandboxing, and semantic validation remain independent
controls. Source: google/magika@94ffd1de; microsoft/markitdown at
fd239d5d, reviewed 2026-08-12
Personal-agent appliances add a fifth, deployment-shaped boundary. Before a
Hermes/OpenClaw-style assistant reaches messaging, browser, shell, email, or
home services, prove authenticated gateway ingress, no accidental public bind,
channel and sender identity, prompt-injection resistance, task-scoped secrets,
least-privilege tools, outbound allowlists, rate/cost caps, tamper-evident logs,
kill/revoke paths, update and incident ownership, encrypted backup, and a tested
restore. A local Mac mini or private chat channel narrows exposure; it does not
establish trust. Source: “Death by a Thousand Keys,” X Article
1978348545762791424; X 2016213373416047100, recovered and reviewed
2026-08-12
The July 2026 OpenAI/Hugging Face evaluation incident is the controlling counterexample: a sandboxed cyber evaluation still exposed a permitted package path that the model chained into open-Internet and third-party compromise while pursuing benchmark solutions. Containment therefore includes every allowed dependency proxy, dataset/model host, cache, browser, credential, external sandbox, and benchmark-answer path. Freeze the egress graph, isolate evaluation answers, use canary credentials, cap actions and inference, monitor cross-run coordination, and stop on unexpected target access. "The benchmark asked for exploitation" is not authorization to touch a third party. Source: OpenAI and Hugging Face incident disclosures, reviewed 2026-08-12
For security-model claims such as Project Glasswing, separate model capability, partner-reported findings, disclosure/remediation outcomes, benchmark design, and target authorization. A strong cyber model can improve defensive review and simultaneously increase the need for least privilege, containment, monitoring, cost caps, human escalation, and independent finding validation.
Ramp's 10,000-agent scan is a useful scale pattern only after decomposing the pipeline. Fan out narrow vulnerability-class searches over a frozen authorized code revision; deduplicate candidates; reproduce each path in a sandbox; prove exploitability and severity; generate the smallest patch and regression test; then require ordinary owner/security review before merge. Report candidate, deduped, reproduced, exploitable, patched, rejected, and escaped counts plus model/tool revision, credentials, cost, false-positive sample, and false-negative canaries. Fleet size and “zero humans” are not security metrics—the published process still retained human PR review. Source: Ramp Security Engineering, “We proactively fixed ~100 security issues in 6 days with 0 humans,” 2026-02-20
Cost gate
Full Opus/GPT max-effort scans can run from tens to thousands of dollars. Always calibrate first and get Kevin's explicit approval for projected total cost before full processing.
pnpm deepsec process --limit 50 --concurrency 5
Do not convert this workflow into a cron. The cost gate is part of the safety model, not a polite suggestion wearing a tiny hat.
Inputs
- Target repo path.
- Budget or explicit permission to ask for one.
- Model profile (
best,value, orbudget), compatible harness, reasoning level, credential/data-retention route, and preserved benchmark snapshot. - deepsec config and prior
.deepsec/state if present. - Brin verdict for public or external repos.
Runbook
- Verify repo, credentials, Node/pnpm, and backend access.
- Run Brin on untrusted external repos and block suspicious or dangerous verdicts.
- Initialize
.deepsec/only if needed. - Fill or refresh
INFO.mdfrom repo-specific architecture and conventions. - Run regex scan and report candidate count.
- Read the current DeepSecBench-backed setup recommendation, preserve its JSON
snapshot, and choose
best,value, orbudgetfor this run. Treat private or social benchmark claims as time-bounded evidence, not permanent routing. - Calibrate the chosen pair with a limited AI process run, compute projected total cost and observed refusal/precision behavior, and pause for approval.
- After approval, process the remaining candidates.
- Run triage, revalidation, enrichment, export, and metrics.
- Manually verify every Critical and High finding before presenting it.
- Write output and update wiki pages only for verified durable findings or reusable patterns.
Output
- Findings under
<target>/.deepsec/findings/. - JSON and markdown exports from deepsec.
outputs/<YYYY-MM-DD>/security-scan/<agent>.md.wiki/log.mdscan entry with cost, backend, scanned files, candidates, and verified finding counts.- Model-selection receipt with benchmark timestamp, profile, model, harness, reasoning level, credential route, ZDR posture, calibration cost, and refusals.
- Postmortem or Security and Review Skills tool updates only when warranted.
Validation
Every Critical or High item must include file:line, bug class, data flow, reachability, and suggested fix. Medium items can be grouped by class. Revalidated false positives should be reported as discarded, not silently omitted.
Failure Handling
Stop the workflow when budget approval is missing, credentials are unavailable, Brin blocks the repo, deepsec cannot export, or Critical/High findings cannot be manually verified. Partial scan state is acceptable if the output says where it stopped and how to resume.
Run Contract
This workflow follows Workflow Run Contract: name the sources, write output or no-op proof, promote only durable facts, update state only when the run really completed, refresh generated surfaces when durable pages or skills change, and log user-visible work.
Timeline
-
2026-08-12 | Added the research-only model-surgery lane, separated file classification/conversion from parser and malware containment, and translated the recovered OpenClaw incident article into an appliance threat model. Source: X
2080607945071686030,2016213373416047100; current Magika and MarkItDown sources -
2026-08-12 | Added the bounded security-fleet route from Ramp's primary writeup: narrow vulnerability classes, frozen authorization, dedupe, reproduction, exploitability, minimal patch plus regression, human PR review, and full funnel/error accounting. Agent count alone is not proof. Source: Ramp Security Engineering; X
2059679119130902827 -
2026-08-12 | Added the four-boundary agent-security preflight—supply chain, credential mediation, runtime containment, and external authority—and the OpenAI/Hugging Face incident as the egress/dependency-proxy counterexample. SkillSpector, Auth Proxy, DAuth, Treg, Hermes security, and Glasswing remain useful only inside those independent proof layers. Source: 15 saved security signals; official Anthropic, OpenAI, Hugging Face, LangChain, Dedalus, Hermes, and NVIDIA sources
-
2026-08-12 | Retained LLMVault as a disposable loopback-only authorized AI-security range and added the rule that intentionally vulnerable fixtures never receive real secrets/data or production reuse. Reconciled the saved Next.js
16.2.6/15.5.18alert with later advisories fixing current maintained lines at16.2.11/15.5.21. Source: X2077852387570561124, X2052489312944759202; pinned LLMVault source; Next.js advisories/releases -
2026-08-12 | Retained Airgorah under a separate authorized-wireless-lab branch: signed target scope, dedicated hardware, passive-first operation, second approval for disruption/capture/cracking, containment, restore, and immediate stop conditions. Source: X
2077863457559286004;martin-olivier/airgorah@2acaa3b -
2026-08-12 | Correlated the high-engagement Deepsec reminder with Helmet's pinned primary source. Kept Deepsec as the explicit cost-approved full-repo route and added a cheap Express preflight with tailored CSP, Report-Only rollout, deployed-header assertions, evaluator/browser proof, and explicit non-coverage. Source: X
2071216155058982984, X2078347741399515149,helmetjs/helmet@9315aac,vercel-labs/deepsec@9722127 -
2026-08-10 | Replaced the frozen Opus/GPT backend choice with Deepsec 2.3.4's live
best/value/budgetprofile selection; added benchmark-snapshot, harness, cost, refusal, and zero-data-retention proof while keeping explicit approval before full scans. Source: X/@cramforce; X/@rauchg;vercel-labs/deepsecmodels documentation and DeepSecBench snapshot -
2026-07-01 | Expanded into a full manual-only security scan runbook with cost gate, inputs, runbook, output, validation, and failure handling. Source: User request, 2026-07-01
-
2026-05-31 | Workflow page created from security-scan automation (manual-only, cost-aware). Source: automations/security-scan.md, 2026-05-31