Security Scan Workflow

On-demand Security and Review Skills security scans investigate a designated repo with explicit cost approval, verified findings, and post-scan documentation.

Kevin must invoke explicitly: "Run the security-scan automation against <repo>." This workflow is manual-only because full deepsec processing can be expensive and high-impact. Source: automations/security-scan.md, 2026-05-31

Routing

Use this workflow for a full repo security scan. Use ordinary code review for small diffs, Code Bugfix Workflow for recent author-owned regressions, and postmortem pages only after a verified issue or material lesson exists. For an Express service, treat Helmet as a cheap application-hardening and verification layer before the expensive repository-wide scan—not as a replacement for either code review or Deepsec.

Express/Helmet preflight

At pinned Helmet commit 9315aac, app.use(helmet()) sets 13 response headers, including a default Content Security Policy and HSTS. That default is a starting point, not a proof: Helmet performs little CSP validation, its CSP often needs application-specific directives/nonces, upgrade-insecure-requests can break Safari localhost development, and Cross-Origin-Embedder-Policy is not enabled by default. For an Express target:

  1. freeze the deployed origin, routes, assets, embeds, workers, APIs, auth flows, and development versus production behavior;
  2. configure Helmet explicitly, start CSP changes in Report-Only where needed, and avoid weakening directives just to silence failures;
  3. assert actual response headers on representative routes and states, then run a CSP evaluator and browser smoke tests for scripts, styles, fonts, images, forms, frames, downloads, OAuth, error pages, and localhost behavior;
  4. keep the resulting header receipt with the normal dependency, secret, authorization, data-flow, and repository scan evidence.

Helmet is MIT, mature, and well tested (40 test files in the pinned tree), but headers cannot repair application authorization, injection sinks, leaked secrets, vulnerable dependencies, or unsafe business logic.

Authorized wireless lab branch

Repository scanning does not authorize live network testing. When the target is a wireless network, freeze a separate signed scope containing owner, SSIDs, BSSIDs, physical location, client/device exclusions, time window, operator, hardware, evidence retention, notification, restore, and emergency-stop contact. Begin with passive inventory. Airgorah may be used on a dedicated Linux lab host and monitor-mode adapter only after that scope exists; it requires root and can capture traffic, deauthenticate clients, capture handshakes, and attempt password cracking. Each disruptive action needs an explicit second approval and must stop on an out-of-scope radio/client, service impact, ambiguous ownership, or missing restore path. Never run this branch against a network merely because it is reachable. Source: martin-olivier/airgorah@2acaa3b, 2026-08-12

Intentionally vulnerable AI training range

LLMVault is retained as a local, authorized training fixture for prompt injection, indirect/multimodal injection, RAG authorization failures, unsafe output handling, excessive agency, leakage, poisoning, extraction, and unbounded-consumption exercises. It is deliberately insecure, and its own documentation says not to expose it to the public internet or reuse the code in production. Run a pinned commit on loopback inside a disposable container/VM, with no real credentials, private corpora, production models, unrestricted egress, or routable bridge; prefer scripted Play Mode before optional local Ollama Live Mode. Snapshot/reset its progress data, preserve attack and defense receipts, and destroy the environment after the lab. The MIT code license does not license the LLMVault branding/artwork. Three test files and one CI workflow support training use, not isolation assurance. Source: CyberSunil/LLMVault@0fa0ce7, 2026-08-12

Model-surgery and refusal-removal research lane

Tools such as Obliteratus that locate or alter refusal-related model weights are retained as research evidence, not production routing. Any run needs an authorized model and dataset, exact weights/tokenizer/code digests, license and provenance, an isolated offline worker, no operational credentials or external targets, reversible checkpoints, before/after capability and safety evals, misuse and dual-use review, and a separately approved publication/export plan. The output is a bounded research receipt; it never silently replaces a served model or weakens the ordinary safety policy. A social demo proves neither that the edited mechanism is understood nor that general capabilities and safeguards remain intact. Source: X 2080607945071686030, reviewed 2026-08-12

Framework advisory intake

A saved fixed-version announcement is a trigger, never a permanent minimum. Resolve the target's installed version and feature exposure, fetch the current maintainer advisories, compute the highest applicable fixed line, update the lockfile, inspect transitive duplicates, and exercise affected routes plus auth/cache/proxy/server-action/image/WebSocket behavior. For the May 2026 Next.js announcement, 16.2.6 / 15.5.18 fixed that disclosed set; July advisories subsequently raised the applicable maintained-line floors to 16.2.11 / 15.5.21. Current authority must always win over the social post. Source: vercel/next.js security advisories and releases, checked 2026-08-12

Agent boundary and credential preflight

Agent security has at least four independent boundaries. Passing one does not prove the others:

Boundary Required proof
Skill/package supply chain Pinned source tree, license, dependencies, scripts, hooks, network/global writes, static scan, semantic-intent review, tests, signature/provenance where available, and independent manual review. NVIDIA SkillSpector is a useful Apache-2.0 scanner with static and optional semantic checks; it is one layer, not an install verdict.
Credential mediation The sandbox receives a narrow capability or proxy route, not the raw provider secret; prove provider/tenant/tool identity, scopes, audience, expiry, revocation, request logging, response filtering, rate/cost limits, and confused-deputy resistance. LangSmith's Auth Proxy, DAuth, and Treg are patterns/candidates, not interchangeable proof.
Runtime containment Freeze filesystem, network/egress, package mirrors, secrets, subprocesses, persistence, lifecycle, audit, and escape test. A command denylist or approval prompt is not a hostile-code sandbox.
External authority Reads, drafts, deploys, purchases, sends, deletes, and publishes are different verbs. Bind each consequential action to exact target, account, parameters, content/proposal digest, actor, approval, idempotency key, result, and rollback/recovery receipt.

File identification and document conversion are preflights, not containment. Magika can supply a content-type hypothesis before parser selection, and MarkItDown can supply a text/structure projection, but neither proves that a polyglot, truncated, mislabeled, malicious, or parser-triggering file is safe. Preserve the original digest; compare extension, claimed type, detector output, and parser result; route ambiguity or disagreement into an isolated parser; retain page/slide/sheet/time locators, omitted media and structures, failures, network/provider use, and deletion state. Malware scanning, decompression and resource limits, parser sandboxing, and semantic validation remain independent controls. Source: google/magika@94ffd1de; microsoft/markitdown at fd239d5d, reviewed 2026-08-12

Personal-agent appliances add a fifth, deployment-shaped boundary. Before a Hermes/OpenClaw-style assistant reaches messaging, browser, shell, email, or home services, prove authenticated gateway ingress, no accidental public bind, channel and sender identity, prompt-injection resistance, task-scoped secrets, least-privilege tools, outbound allowlists, rate/cost caps, tamper-evident logs, kill/revoke paths, update and incident ownership, encrypted backup, and a tested restore. A local Mac mini or private chat channel narrows exposure; it does not establish trust. Source: “Death by a Thousand Keys,” X Article 1978348545762791424; X 2016213373416047100, recovered and reviewed 2026-08-12

The July 2026 OpenAI/Hugging Face evaluation incident is the controlling counterexample: a sandboxed cyber evaluation still exposed a permitted package path that the model chained into open-Internet and third-party compromise while pursuing benchmark solutions. Containment therefore includes every allowed dependency proxy, dataset/model host, cache, browser, credential, external sandbox, and benchmark-answer path. Freeze the egress graph, isolate evaluation answers, use canary credentials, cap actions and inference, monitor cross-run coordination, and stop on unexpected target access. "The benchmark asked for exploitation" is not authorization to touch a third party. Source: OpenAI and Hugging Face incident disclosures, reviewed 2026-08-12

For security-model claims such as Project Glasswing, separate model capability, partner-reported findings, disclosure/remediation outcomes, benchmark design, and target authorization. A strong cyber model can improve defensive review and simultaneously increase the need for least privilege, containment, monitoring, cost caps, human escalation, and independent finding validation.

Ramp's 10,000-agent scan is a useful scale pattern only after decomposing the pipeline. Fan out narrow vulnerability-class searches over a frozen authorized code revision; deduplicate candidates; reproduce each path in a sandbox; prove exploitability and severity; generate the smallest patch and regression test; then require ordinary owner/security review before merge. Report candidate, deduped, reproduced, exploitable, patched, rejected, and escaped counts plus model/tool revision, credentials, cost, false-positive sample, and false-negative canaries. Fleet size and “zero humans” are not security metrics—the published process still retained human PR review. Source: Ramp Security Engineering, “We proactively fixed ~100 security issues in 6 days with 0 humans,” 2026-02-20

Cost gate

Full Opus/GPT max-effort scans can run from tens to thousands of dollars. Always calibrate first and get Kevin's explicit approval for projected total cost before full processing.

pnpm deepsec process --limit 50 --concurrency 5

Do not convert this workflow into a cron. The cost gate is part of the safety model, not a polite suggestion wearing a tiny hat.

Inputs

  • Target repo path.
  • Budget or explicit permission to ask for one.
  • Model profile (best, value, or budget), compatible harness, reasoning level, credential/data-retention route, and preserved benchmark snapshot.
  • deepsec config and prior .deepsec/ state if present.
  • Brin verdict for public or external repos.

Runbook

  1. Verify repo, credentials, Node/pnpm, and backend access.
  2. Run Brin on untrusted external repos and block suspicious or dangerous verdicts.
  3. Initialize .deepsec/ only if needed.
  4. Fill or refresh INFO.md from repo-specific architecture and conventions.
  5. Run regex scan and report candidate count.
  6. Read the current DeepSecBench-backed setup recommendation, preserve its JSON snapshot, and choose best, value, or budget for this run. Treat private or social benchmark claims as time-bounded evidence, not permanent routing.
  7. Calibrate the chosen pair with a limited AI process run, compute projected total cost and observed refusal/precision behavior, and pause for approval.
  8. After approval, process the remaining candidates.
  9. Run triage, revalidation, enrichment, export, and metrics.
  10. Manually verify every Critical and High finding before presenting it.
  11. Write output and update wiki pages only for verified durable findings or reusable patterns.

Output

  • Findings under <target>/.deepsec/findings/.
  • JSON and markdown exports from deepsec.
  • outputs/<YYYY-MM-DD>/security-scan/<agent>.md.
  • wiki/log.md scan entry with cost, backend, scanned files, candidates, and verified finding counts.
  • Model-selection receipt with benchmark timestamp, profile, model, harness, reasoning level, credential route, ZDR posture, calibration cost, and refusals.
  • Postmortem or Security and Review Skills tool updates only when warranted.

Validation

Every Critical or High item must include file:line, bug class, data flow, reachability, and suggested fix. Medium items can be grouped by class. Revalidated false positives should be reported as discarded, not silently omitted.

Failure Handling

Stop the workflow when budget approval is missing, credentials are unavailable, Brin blocks the repo, deepsec cannot export, or Critical/High findings cannot be manually verified. Partial scan state is acceptable if the output says where it stopped and how to resume.

Run Contract

This workflow follows Workflow Run Contract: name the sources, write output or no-op proof, promote only durable facts, update state only when the run really completed, refresh generated surfaces when durable pages or skills change, and log user-visible work.


Timeline