Cloudflare security-audit-skill Breakdown: Six-Phase Orchestration Where the Verifying Agent Never Found the Bug

1. At a Glance: Is It Worth Your Time

Rating: ★★★★☆ (4 / 5)

Let’s start with this project’s first anomaly: 22,616 stars, only 14 commits.

I dug through the commit history. Of those 14 commits, only two are real reworks (“Rework the audit workflow, findings contract, and validators end to end” on 2026-09-10, and “Bring the whole skill into one coherent house style” on 2026-07-06). The repo has 21 files in total, 14 of which are Markdown prompt documents; the only executable code is 4 .cjs files (two validators + two tests) and 1 JSON schema.

So this is not “software” — it is a machine-readable version of an audit methodology. The stars come mainly from the Cloudflare brand and the Build your own vulnerability harness blog post — the README states plainly that this skill is the “single-repo starting point” of that harness, which later grew into a fleet-wide multi-stage system.

Why does it still deserve 4 stars? Because it attacks the most real pain point in AI auditing: agents lie.

Put a coding agent on a code review and you get a pile of “found SQL injection” conclusions; click into them and the function it cites does not exist, or the entry point it describes is not externally reachable at all. The reason is simple: the “enthusiasm” an agent has when generating conclusions and the “rigor” it has when generating evidence are not the same mechanism. Cloudflare’s fix is refreshingly plain — split the work across different agents:

  • The agent that finds a vulnerability (the hunter) is not allowed to confirm it;
  • Every candidate goes to a fresh verifier whose job is not to “validate” but to disprove;
  • A final round of record verification sends yet another batch of fresh agents to check whether “the source code this finding cites actually says that”;
  • If that step produced a material replacement, a further independent verifier reviews the replacement itself.

For a conclusion to survive into REPORT.md, it must pass through three batches of mutually unaware agents. In prompt-engineering terms, that is a big deal.

The second counterintuitive move is daring to write this design principle down:

Multiple runs improve coverage. In our test runs, a single run found roughly half of the vulnerabilities that repeated runs found in total.

“A single run finds about half of what repeated runs find” — the vendor says outright that one run misses roughly 50%. This is the same kind of honesty as the Cisco Skill Scanner piece (Issue 016, which dared to publish a 7.75% recall). The pattern I keep seeing in this space: everyone races on “how many did I find”, nobody says “how many did I miss”. Cloudflare at least gives an order of magnitude.

The deductions are equally clear: 14 commits / 3 months of history — engineering maturity nowhere near the star count; no quantitative evaluation at all — no confusion matrix like Cisco’s, and the “half” figure comes from a design principle, not a benchmark; the sandbox requirement is very hard — without an OS-enforced sandbox it deliberately downgrades every lead to needs_validation and refuses to execute target code, so many people will finish a run and find “nothing was confirmed”; and cost — 6 phases × multiple hunters × officially recommended multiple runs means the token bill for one full audit is not small.

Key Data (as of 2026-10-07)

Vendor Cloudflare (security-ai-research@cloudflare.com)
Type Coding-agent skill (prompt orchestration, not standalone software)
Language JavaScript (only 4 zero-dependency .cjs validators)
Stars / Forks 22,616 / 1,315
Commits / Open Issues 14 / 51
License MIT
First commit 2026-06-18 (~3 months)
Last commit 2026-09-14 (15 days ago)
File composition 21 files: 14 Markdown prompts + 4 .cjs + 2 tests + 1 schema
Attack-class docs 12 (Core / Memory-safety / AI-LLM / Web-protocol / Client-side / Supply-chain / Cloud / RPC / Resource / Data-isolation / Desktop-mobile)
Verdict types confirmed / needs_validation / rejected
Install npx skills add https://github.com/cloudflare/security-audit-skill --skill security-audit

Who It’s For

Audience Score Why
Developers with a coding agent 5 / 5 One security audit this codebase and it runs — zero configuration, and the methodology is free to take
Security teams / red team 4 / 5 The 12 attack-class docs are an excellent checklist on their own and can be extracted for standalone use
People researching AI audit methods 5 / 5 Rare public material that turns “how to stop agents from deceiving themselves” into an executable workflow
Teams needing compliance deliverables 2 / 5 Output is findings.json + three Markdown files, not SARIF — it will not feed your existing vulnerability management platform
CI gating 2 / 5 Long multi-agent workflow + high cost + no stable baseline — not something to trigger on every commit
Anyone expecting a “vulnerability scanner” 1 / 5 It is audit orchestration, not a scanner; you will be disappointed

Building on It

Difficulty Notes
Editing attack classes / adding domains ★☆☆☆☆
Reusing the validators ★☆☆☆☆
Wiring into your own reporting system ★★★☆☆
Changing the workflow orchestration ★★★★☆

2. What It Is, What It Isn’t

It is:

  • A set of Markdown agent-orchestration specifications that make your coding agent perform a security audit in 6 phases
  • An attack-class knowledge base (12 documents, covering everything from memory safety to LLM prompt injection)
  • Two zero-dependency validators that gate the structure of findings.json and coverage-ledger.json
  • The public, simplified version of Cloudflare’s internal vulnerability-discovery harness

It is not:

  • Not a scanner. It does not grep, does not run SAST, does not do taint analysis. It directs your agent to do those things. If you want a scanner that emits SARIF, see Issue 012 Vigolium (Go-native, 324 modules) or Issue 016 Cisco Skill Scanner.
  • Not a vulnerability database. What it produces is a record of this one audit, not a CVE library.
  • Not deterministic. Run it twice on the same repo and the results differ. This is by design (it stacks coverage across multiple runs), and it also means it cannot serve as a CI gate.
  • No guarantee of completeness. The vendor’s own words: a single run finds roughly half.

One easily misunderstood point deserves its own paragraph: what it scans is not “skills” — it is any codebase. The name contains “skill” because it is itself an agent skill. It is a completely different category from Issue 015 SkillSpector (NVIDIA’s scan of 31k skills) and Issue 016 Cisco Skill Scanner. Those two scan other people’s skills for malice; this one performs a security audit of your own repository.

3. Core Capabilities

3.1 Three Batches of Mutually Unaware Agents (Adversarial Validation)

This is the foundation of the entire design. The README states it bluntly:

Adversarial validation. The agent that checks a finding is never the agent that found it.

In workflow terms there are four gates:

  1. The hunter (Phase 2) takes a task from a ledger unit, does the work, and records what it checked;
  2. The verifier (Phase 3) receives a brand-new context, and its task is to disprove (tries to disprove it), not to restate;
  3. The record verifier (Phase 5) is yet another batch of fresh agents, dedicated to checking whether the source-code claims cited in the final record are actually true;
  4. If step 3 produced a material replacement, another independent verifier reviews that replacement.

Item 4 is the one I find most ruthless — even the “correction” itself is not trusted.

3.2 The Coverage Ledger (coverage-ledger.json)

This is its most fundamental difference from traditional SAST: organized by “coverage”, not by “rules”.

Phase 1 reconnaissance produces architecture.md + coverage-ledger.json. The ledger slices the codebase into audit units; each hunter claims units, does the work, and records what it checked. Dedicated coverage critics then hunt for holes — which units nobody claimed, which claims of “checked” have no corresponding evidence in the record.

The parent process runs validate-coverage-ledger.cjs right after the ledger is created, and again after every subsequent ledger update. This “validate on every update” rule is critical: it stops the ledger from being quietly corrupted over a long workflow.

The upside is that multiple runs are additive:

Multiple runs against the same repo are additive. The skill uses prior ledgers and findings to target gaps, revalidate changed source, and carry forward current-source evidence without treating stale or unresolved work as covered.

That last clause is the key — stale and unresolved work is never counted as “covered”. Many incremental scanning tools get this wrong: files scanned last time are skipped this time, so last time’s needs_validation never gets reprocessed.

3.3 Three Verdicts, Strictly Separated Semantics

Verdict Meaning Has severity?
confirmed Complete source trace and bounded observed result Yes
needs_validation Carries one exact unresolved fact No
rejected A disproven candidate, still kept on record No

I strongly agree with the design of giving needs_validation no severity — why would something unvalidated get a rating? It forces you to treat it as a to-do item rather than a “low-severity vulnerability”, avoiding the classic failure mode of “a flood of INFO-level alerts drowning the real problems”.

Keeping rejected on record instead of discarding it is also worth copying: a disproven path is itself audit evidence (it tells the next hunter this road was walked — don’t walk it again).

3.4 Several “Counterintuitive but Correct” Judgment Principles

These are worth more, in my view, than the workflow itself:

  • Only confirm established boundary failures. If it does not rise to that bar, put it in needs_validation and write down the exact unresolved fact.
  • Severity requires impact. Severity = likelihood × impact, not “how far it deviates from a checklist”. This one aims squarely at the black magic of CVSS scoring.
  • Defense-in-depth gaps are not vulnerabilities. If layer A already blocks the attack, a missing layer B is a hardening suggestion, not a vulnerability. This rule alone could kill half of all invalid tickets.
  • Target-neutral reporting. The report stays neutral toward the target — no going soft because you are auditing your own project.

4. Architecture: What the Six Phases Do

Phase What it does Artifacts
1. Reconnaissance Map architecture, trust boundaries, input surface, existing evidence, deterministic coverage architecture.md + coverage-ledger.json
2. Coverage-led hunting Dispatch hunters from ledger units, record checks, coverage critics hunt for holes candidate findings
3. Candidate validation Each unique candidate goes to a fresh verifier to be disproved validated candidates
4. Structured output Write confirmed / needs_validation / rejected into findings.json, validated against report-schema.json findings.json (run validate-findings.cjs)
5. Independent record verification Fresh agents check the final source claims; material replacements get another independent verifier verified records (validators re-run after every replacement)
6. Target-neutral reporting Derive the report from verified records + the coverage ledger REPORT.md + FINDINGS-DETAIL.md + NEEDS-VALIDATION.md

The validator invocation timing is deliberate:

  • validate-coverage-ledger.cjs — run immediately after Phase 1 creates the ledger, then after every ledger update (Phases 1–5)
  • validate-findings.cjs — run once in Phase 4, and again after every replacement in Phase 5

This “re-validate after every change” pattern is the standard answer for long-workflow agent orchestration. Six phases can take tens of minutes; intermediate artifacts drift easily, and a single end-of-pipeline validation cannot catch it.

The 12 Attack-Class Documents

This is the real content asset, and the only part usable outside the workflow:

Document Coverage
ATTACK-CLASSES.md Core / wildcard / obvious-things
MEMORY-SAFETY-AND-BINARY.md Memory safety, binaries, kernels (native targets)
AI-AND-LLM.md Prompt injection, agents/tools, output handling (LLM targets)
WEB-PROTOCOL-AND-AUTH.md HTTP request framing, caching, authentication protocols
CLIENT-SIDE.md DOM injection, message trust, UI redirection, prototype pollution
SUPPLY-CHAIN-AND-RELEASE.md Dependencies, CI, releases, signing, updates, plugins, extensions
CLOUD-AND-DEPLOYMENT.md IAM, IaC, containers, serverless, ingress, runtime configuration
PROTOCOLS-RPC-AND-MESSAGING.md RPC, serialization, queues, brokers, webhooks, streaming protocols
RESOURCE-EXHAUSTION-AND-AVAILABILITY.md Shared resources, quotas, queues, workers, operator spend
DATA-ISOLATION-AND-LIFECYCLE.md Tenant isolation, caching, search, exports, backups, migrations, deletion, recovery
DESKTOP-MOBILE-AND-LOCAL-IPC.md Native apps, deeplinks, webviews, exported components, helpers, daemons, local IPC

RESOURCE-EXHAUSTION explicitly lists operator spend as a category — in the AI agent era this is a real attack surface (make your agent burn tokens frantically / hammer paid APIs), and it is rarely seen in traditional checklists.

5. Ecosystem and Benchmarks

This must be said clearly: there is no benchmark.

Cisco Skill Scanner (Issue 016) published an evaluation with a confusion matrix; SkillSpector (Issue 015) scanned 31k skills and produced statistics. This Cloudflare skill has not a single number — the only quotable quantitative statement is that line in the design principles: “a single run found roughly half of the vulnerabilities that repeated runs found in total”.

That sentence’s provenance is “In our test runs”, not a public benchmark. I treat it as an order-of-magnitude reference, not a reproducible metric.

It does have a few other “credibility sources”:

  • Lineage: the README states plainly that this is the starting point of Cloudflare’s internal vulnerability-discovery harness, which has grown into a fleet-wide system. That is a “we really run this in production” claim — unverifiable, but meaningful.
  • 52 open issues (an extremely low ratio against 22.6k stars — meaning either few users or few problems; both are plausible).
  • 1,315 forks — plenty of people modifying it, and for a set of prompts, forking is the primary mode of “use”.

Where it sits against peers:

Tool Category Essence Output
Cloudflare security-audit-skill Audit orchestration Directs your agent to audit findings.json + 3 Markdown files
Cisco Skill Scanner (016) Skill scanner 9 engines, static + LLM SARIF / HTML / JSON
SkillSpector (015) Skill scanner NVIDIA large-scale scanning statistics + reports
deepsec (011) One-click audit Vercel’s wrapped full audit report
LuaN1ao (009) Pentest reasoning chain Pins conclusions back to evidence evidence chain

Its unique position: the only public artifact whose core design goal is “stop the agent from deceiving itself”. Every other tool improves “discovery capability”; this one improves “conclusion trustworthiness”.

6. Deployment

Install

1
2
3
4
5
6
7
8
# Project-level
npx skills add https://github.com/cloudflare/security-audit-skill \
--skill security-audit

# User-level
npx skills add https://github.com/cloudflare/security-audit-skill \
--skill security-audit \
--global

Launch your coding agent inside the target repo (or pointed at it), then simply say:

1
2
3
security audit this codebase
find security vulnerabilities in ./src
do a security review, output to ~/audits/my-project

The skill activates automatically when the request matches its trigger phrases (security audit / find vulnerabilities / pen-test the code, etc.).

Two Modes

  • Full audit mode — a direct “audit this codebase / do a pen-test” request. Produces the complete report artifacts.
  • Guidance mode — security questions, focused vulnerability analysis. This is the default, unless you explicitly ask for the report artifacts.

With no output directory specified, the default is ~/security-audit-skill/<repo-name>/run-<N>. Note that run-<N> — it exists precisely for “additive multiple runs”: the second run opens run-2 and reads run-1‘s ledger.

Hard Prerequisites (the Easiest Thing to Trip On)

  1. A coding agent that supports tool calls + parallel sub-agents (Claude Code / Codex class; agents without parallel sub-agents cannot run multiple hunters)
  2. Node.js — only to run the two zero-dependency validators
  3. An OS-enforced sandbox for executing target-controlled code (build / test / process / browser / emulator / fuzzer / fixture). Requirements:
    • external networking disabled
    • a sanitized, allowlisted environment
    • enforced resource limits
    • writes allowed only to designated scratch paths

Item 3 is not a suggestion — it is a switch. The README, verbatim:

Without these controls, the workflow keeps the lead as needs_validation instead of executing target code.

No sandbox → every lead stalls at needs_validation → you get a report that “confirms nothing”. That is a safety design (never run malicious code bare during an audit), but if your sandbox is not set up, the experience is simply “this tool is useless”.

On “Not Writing Back into the Repo”

The README has one very restrained sentence:

The workflow writes inside the target repository only when you explicitly select a directory that version control ignores.

By default all artifacts live under ~/security-audit-skill/ — your git working tree is not polluted. Compared with agent tools that stuff an .audit/ directory into your repo, this default is decent behavior.

7. Boundaries and Risks

1. 14 commits / 3 months of history — do not be fooled by the 22k stars.
The stars come from the Cloudflare brand + that blog post, not from code maturity. Commits are concentrated in two people (literally-dan, Dan Jones). This is a high-quality methodology draft, not a hardened product.

2. Zero quantitative evaluation.
“Half per single run” comes from one “in our test runs” inside a design principle. No test set, no confusion matrix, no reproducible script. You cannot assess how much of your codebase it actually covers. On this axis it is clearly less honest than Cisco Skill Scanner, which turned its ugly number (7.75% recall) into a dedicated section.

3. Cost is a real barrier.
6 phases × multiple hunters × coverage critics × three rounds of validation × officially recommended multiple runs. The token bill for one full audit may be an order of magnitude higher than you expect, and “run it several more times to approach ~100% coverage” is the vendor’s own advice. On large repos, start small.

4. No sandbox = wasted run.
See the previous section. This is the most common cause of “installed it and it does nothing”.

5. It does not produce compliance deliverables.
findings.json uses a custom schema, not SARIF. It will not feed DefectDojo / Jira security tickets or similar existing pipelines — you have to write the mapping yourself.

6. Non-determinism.
Same input, two runs, different results. Never use it as a CI gate — you will collect a pile of flaky build failures. Its positioning is “periodic, human-triggered deep audit”, not “a checkpoint on every commit”.

7. The prompt-injection surface.
Auditing your own repo is fine; if you point it at third-party / untrusted codebases, the comments and docs inside the code are themselves injection carriers (AI-AND-LLM.md explicitly lists the prompt-injection class). The official sandbox requirements block part of the execution surface, but the text the agent reads cannot be blocked.

8. Your concern may already be among the 51 open issues.
The ratio against 22.6k stars is low, but the project is only 3 months old — issue accumulation has just begun.

8. Getting Started (in This Order)

  1. Read only SKILL.md and VALIDATION-AND-REPORTING.md first. The former has an “audit anti-patterns” section; the latter covers Phases 3–6. After those two you know what it is trying to do — faster than running it blind.
  2. Extract ATTACK-CLASSES.md for standalone use. Even if you never adopt this workflow, that attack-class list is worth having as a manual code-review checklist. Auditing an LLM app? Read AI-AND-LLM.md. Auditing web? Read WEB-PROTOCOL-AND-AUTH.md.
  3. Set up the sandbox before your first run. External networking disabled + resource limits + scratch-path-only writes. Cut corners here and everything downstream is needs_validation.
  4. Use a small directory for the first run, not the whole repo. Try path-scoped instructions like find security vulnerabilities in ./src first to gauge token consumption.
  5. Confirm your agent supports parallel sub-agents. Without it, multiple hunters degrade to serial execution and it takes forever.
  6. Specify the output directory explicitly (do a security review, output to ~/audits/my-project) instead of letting it pile up default run-<N> folders.
  7. Run it at least twice before drawing conclusions. The vendor says one run finds about half. The second run reads the first run’s ledger to fill the holes — that is the moment this design actually pays off.
  8. Read NEEDS-VALIDATION.md first, not REPORT.md. That list of “exact unresolved facts” often tells you more about where the workflow gets stuck on your code than the confirmed items do.
  9. To integrate with your own platform, start from report-schema.json. It is standard JSON Schema; mapping it to your finding format is not hard.
  10. Do not use it as a gate. Periodic, human-triggered deep audit — not a CI checkpoint.

9. The One-Line Verdict

Cloudflare’s security-audit-skill earns its 22k stars not through maturity (14 commits, zero quantitative evaluation, 3 months old) but because it is the only public artifact I have seen that turns “how to stop an agent from lying about its own conclusions” into an executable workflow — three batches of mutually unaware agents, disproven paths still recorded, stale work never counted as covered, and a single run finding only half: take any one of these four and your existing AI audit practice improves.

First move for developers: npx skills add https://github.com/cloudflare/security-audit-skill --skill security-audit --global → start using ATTACK-CLASSES.md as a checklist → set up the sandbox → run the first pass on a small module under ./src, and read NEEDS-VALIDATION.md first. Even if you end up not adopting this workflow, the 6-phase design and the 12 attack-class documents are worth a read — its best part is not the code; it is the judgment packed into those 14 Markdown files.

评论Comments