deepsec Breakdown: Vercel Compresses a Full-Codebase Deep Audit into One npm Command — and the Bill Can Reach Five Figures
1. At a Glance: Is It Worth Your Time
Rating: ★★★★☆ (4 / 5)
Why this score: it has thought its positioning through clearly, and dares to put the price tag on its face. Almost every AI security tool competes on “fast” and “cheap” — plug into CI, scan diffs, results in minutes. deepsec goes the opposite way: it states plainly that it is on-demand review of all code in existing large-scale repos, targeting “hard-to-find problems that have lurked in an application for a long time,” and it defaults to the best models at the maximum thinking level.
What I appreciate most is that the very first paragraph of the README spells out the cost:
…meaning scans can cost thousands or even tens-of-thousands of dollars for large codebases. Our customers have found the cost worth it for how quickly they were able to patch vulnerabilities that would have otherwise gone unfixed.
In a field that is broadly vague about token costs, that candor is scarce. The companion flags --max-cost-usd 100 --max-duration 2h also make the cost controllable — you can nail down the ceiling before the run starts.
The missing star comes down to: cost is a real barrier — individuals and small teams basically cannot afford full mode; the project is only 5 months old (created 2026-04-30), and Vercel Labs projects carry an experimental character by definition; the default path goes through the Vercel AI Gateway — there is a --model-auth direct bypass for bring-your-own keys, but the main path is tied to the Vercel ecosystem; and 69 open issues is not a small number for an 8k-star project.
Key Data (as of 2026-10-01)
| From | Vercel Labs |
| Language | TypeScript (npm package deepsec) |
| Stars / Forks | 8,044 / 488 |
| Commits / Open Issues | 104 / 69 |
| License | Apache-2.0 |
| First commit | 2026-04-30 (~5 months) |
| Last commit | 2026-09-23 (8 days ago) |
| Cost scale | thousands to tens of thousands of dollars for large codebases (officially stated; a cap can be set) |
| Model access | Vercel AI Gateway by default; bring-your-own OpenAI / Anthropic / custom HTTPS provider supported |
| Official site | https://deepsec.sh/ |
Who It’s For
| Audience | Score | Why |
|---|---|---|
| Mid-to-large teams with long-lived repos | 5 / 5 | The positioning fits exactly: one deep full-codebase dig for bugs buried for years |
| Enterprise security operations | 4 / 5 | Cost caps and a sandbox story make it controllable — but you have to win the budget argument first |
| DevSecOps engineers | 3 / 5 | process --diff can gate PRs, but that is not its main scenario |
| Security researchers | 3 / 5 | High engineering completeness, but nothing especially new methodologically |
| Individual developers / small teams | 1 / 5 | Cost alone is a turn-off — full mode was not designed for this scale |
Building on It
| Difficulty | Notes |
|---|---|
| Configuration | Low — npx deepsec init walks you through it interactively: pick a model (with benchmark scores and price comparisons) and a payment method; --max-cost-usd / --max-duration cap it directly |
| Integration | Low — there is report / export --format md-dir / metrics / triage, and the docs cover both CI and coding-agent workflows |
| Kernel | Medium — docs/writing-matchers.md covers writing matchers, there is a plugin mechanism, and Apache-2.0 is friendly |
2. What It Is, What It Isn’t
- Not a fast gate that runs on every commit in CI. Its positioning is on-demand, all-code — not incremental. Diff mode exists via
process --diff, but that is a side mode. - Not a rule-based scanner. The
scanstep does use regex matchers and is free and AI-less, but the real value lives inprocess(the AI investigation step). - Not a replacement for traditional SAST. SAST matches known rule patterns; deepsec has the model reason inside the code context about “is there a problem here.”
- Not limited to the code you wrote.
enrichadds git committer information, and with a plugin it can add ownership data too — meaning it cares about “who should fix this hole.”
Three keywords: all-code, on-demand, maximum thinking level, resumable.
3. The Workflow: scan → process → revalidate
This is the core of understanding deepsec. One audit is split into steps, each with a different cost and value profile:
| Command | Purpose | Uses AI? |
|---|---|---|
scan |
find candidate sites with regex matchers | No (fast, free) |
process |
AI investigates candidates, produces findings + fix suggestions | Yes (the main cost) |
process --diff |
PR mode: scan only files changed in the diff | Yes |
triage |
lightweight P0/P1/P2 triage | Yes (cheap models) |
revalidate |
re-check existing findings, and consult git history to see whether they were already fixed | Yes |
enrich |
add git committer info + (plugin) ownership data | — |
report |
per-project Markdown + JSON summary | — |
export |
JSON per finding, or a directory of markdown files | — |
metrics |
cross-project stats: severity distribution, vulnerability counts by type, true-positive counts | — |
status |
project image snapshot | — |
sandbox <cmd> |
run any of the above on Vercel Sandbox microVMs | Yes |
Design choices I think are well made:
1) scan and process are separated. Free regex matchers narrow the candidates first, so the expensive model only looks at those. That is the most basic respect for token cost — throwing an entire repo at the model directly is not viable.
2) triage uses cheap models for classification. A coarse P0/P1/P2 split does not need the strongest model, and the money saved goes into process.
3) revalidate consults git history. It checks whether existing findings have already been fixed. This is very practical — the biggest problem with a standing audit report is that findings quietly go stale after the first run and nobody maintains them.
4) It is resumable. Ctrl-C, network loss, quota exhausted — rerun the same command and it skips already-analyzed files and picks up the rest. For full audits that run for hours, this is a necessity, not a bonus.
4. Cost: The First Question in Any Evaluation
The official wording is blunt: best models at the maximum thinking level (--thinking-level is adjustable), thousands to tens of thousands of dollars for large codebases.
How should you read that number? My take:
- Its pricing logic is not “a few dollars per commit” but “a deep physical exam a few times a year.” Under that mental model, a few thousand dollars for a batch of vulnerabilities that “would otherwise never get fixed” is a good trade for mid-to-large teams.
- The cost can be nailed down.
--max-cost-usd 100 --max-duration 2hcaps both spend and duration. Setting both before a run is the only responsible way to do it. scanis free. Runscanfirst to see how many candidates there are, use that to estimate the scale ofprocess, and then decide whether to open the budget. That is a good cost-control path.triageuses cheap models, which also holds costs down.
Advice for decision makers: don’t ask “how much does deepsec cost” — ask “how many candidate sites does our repo have, and how much are we willing to pay for one full audit.” Do a small-scale trial with --max-cost-usd, get the real unit price, then extrapolate.
5. Distributed Execution: Fanning the Work Out to microVMs
A large repo is too slow on a single machine, so deepsec can fan out to Vercel Sandbox microVMs:
1 | pnpm deepsec sandbox process --project-id my-app --sandboxes 10 --concurrency 4 |
Two security details deserve their own mention:
- The local working directory is packaged and uploaded, but
.gitis excluded. - Model credentials stay on the host side and are injected only at the designated egress host.
Together, these mean a worker sandbox can get neither the full git history nor the API key.
deepsec’s Own Security Model (official wording, quoted verbatim)
Treat
deepseclike a coding agent with full shell access on the environment that it is running on.
That is, treat it by default as a coding agent with full shell privileges. It is designed to run on trusted input (your source code), but the docs concede: external dependencies or vendored code can carry prompt injection.
Running in sandboxes significantly reduces the exposure:
- the coding agent’s API key is injected outside the sandbox, so it cannot be exfiltrated
- worker sandbox network egress is restricted to the coding agent’s host (egress is allowed during bootstrap, but that phase does not run the coding agent)
Practical conclusion: do not run deepsec bare on a production machine that carries sensitive credentials. Either use sandbox mode, or give it a clean, disposable environment.
6. Quickstart
Initialize
From the root of the repo you want to scan:
1 | npx deepsec init |
This command walks you through the whole flow: pick an AI model (with benchmark scores and price comparisons), pick a payment method (your own OpenAI/Anthropic key, or the Vercel AI Gateway), then it runs unattended — learning the codebase, scanning, AI review.
It adds exactly one .deepsec/ directory to your repo, holding all state and findings.
Cap the Cost
1 | npx deepsec init --max-cost-usd 100 --max-duration 2h |
Get the Report
1 | cd .deepsec |
Follow-Up Scans (run inside .deepsec/)
1 | pnpm deepsec scan # fast pattern scan, free |
Bring Your Own Key (no Vercel account needed)
1 | npx deepsec init --model-auth direct --ai-provider anthropic --ai-api-key-env ANTHROPIC_API_KEY |
The docs are explicit: deepsec stores only the name of the environment variable that holds the key — never the key itself.
If process or revalidate is interrupted because upstream quota or credits ran out, it stops gracefully and tells you where to top up; rerun the same command to resume.
Docs Written for Agents
After initialization, an agent can read documentation that exactly matches the installed CLI version:
.deepsec/node_modules/deepsec/SKILL.md.deepsec/node_modules/deepsec/dist/docs/
When setup fails, these paths are surfaced in absolute, machine-readable form. A thoughtful design — let the coding agent read the docs itself instead of relying on stale APIs from model memory.
7. Boundaries and Risks (the part that must be said honestly)
1) Cost is the hardest barrier, bar none. Thousands to tens of thousands of dollars is not marketing copy — it is the official documentation. Individuals, small teams, and open-source projects are effectively priced out of full mode. Always set --max-cost-usd first.
2) The project is only 5 months old. Created 2026-04-30, 104 commits. For a tool meant to run for hours and spend thousands of dollars, maturity is a real risk. And Vercel Labs projects are experimental in character by definition — the roadmap can change.
3) The default path is tied to the Vercel ecosystem. The default is the Vercel AI Gateway, and distributed execution depends on Vercel Sandbox. --model-auth direct bypasses this entirely, but the main path and the best-practice docs are written around Vercel. If you don’t want to enter that ecosystem, usability takes a discount.
4) Treat it as an agent with full shell access. The official docs say so themselves. External dependencies and vendored code present a prompt-injection surface. Do not run it bare on a machine carrying production credentials.
5) 69 open issues. For an 8k-star, 5-month-old project, that number says the iteration pressure is real. And since each run is expensive, the cost of hitting a bug is high too.
6) Full-mode runtimes are incompatible with CI. That is by positioning, not a defect — but it means you need two tools: something fast in CI, and deepsec for the periodic deep physical. Don’t expect one tool to do both jobs.
7) revalidate is “optional.” The README marks it optional and says it lowers the false-positive rate. Given that AI-audit false positives are not low to begin with, this step should not be skipped if the budget allows — which also means the cost goes up again.
8. Getting Started (in This Order)
- Run
scanalone first. It is free, AI-less, and tells you how many candidates exist. Use that number to estimate the scale ofprocessbefore deciding the budget. Do not skip this step. - Always pass
--max-cost-usdon the firstinit. Start from a number you can afford to lose entirely (say, $50–100), see the real unit price and output quality, then scale up. - Run it first on a mid-sized repo you know well. Familiarity means you can judge how many of its findings are true positives — only you can verify an AI audit’s false-positive rate on the ground. Do not spend big money on an unfamiliar giant repo right away.
- Use
triageto classify andrevalidateto cut false positives. The former saves money downstream; the latter raises the credibility of conclusions. On a tight budget, protectrevalidatefirst. - For sensitive repos, always use sandbox mode.
pnpm deepsec sandbox process --sandboxes 10 --concurrency 4— credentials stay on the host,.gitis not uploaded. Never run it bare on a machine with production credentials. - For CI, use
process --diff. That is the PR-gate mode: changed files only, cost under control. Never put full mode on every commit. - Treat it as “a physical exam a few times a year,” not “a daily gate.” Get that mental model right and both the budget negotiation and the scheduling get much easier.
9. The One-Line Verdict
If you have a repo that has run for years and has never been seriously combed by AI, deepsec is currently the most engineering-complete open-source choice for a full-codebase deep audit — but set --max-cost-usd first, and accept that it is a physical exam, not a gate.
First move for teams: run npx deepsec scan (free) to see the candidate volume → set a budget from that → trial once with npx deepsec init --max-cost-usd 100 → export with export --format md-dir and verify the true-positive rate yourself. If the true-positive rate holds up, then talk about scaling.