SkillSpector Breakdown: NVIDIA Scanned 31k Agent Skills — 26.1% Contain Vulnerabilities, 5.2% Look Likely Malicious

1. At a Glance: Is It Worth Your Time

Rating: ★★★★☆ (4 / 5)

Let me start with why this tool deserves its own article. Agent Skills are the most underestimated attack surface of 2026.

Think about what a skill is: a SKILL.md (natural-language instructions) plus some scripts, and once installed, the agent executes it with implicit trust. It has no lockfile and audit ecosystem like npm packages, and no scanning infrastructure like Docker images — it sits between “document” and “code,” and neither side’s security habits have landed on it.

The numbers SkillSpector gives in its Overview are blunt:

In the 31,132-skill analyzed subset of the research dataset, 26.1% of skills contain vulnerabilities and 5.2% show likely malicious intent.

One in four skills has a vulnerability; one in twenty looks malicious. At those odds, “casually installing a skill” is itself a high-risk action.

Now the tool itself. It is no toy: 71 patterns across 17 categories, covering everything from prompt injection to behavioral AST analysis to taint tracking to YARA malware signatures — including MCP least privilege and tool poisoning. And it is part of the NVIDIA Verified Skills pipeline — NVIDIA itself scans, evaluates, and signs skills with it before publishing its skills catalog. There is a real production pipeline behind the badge.

The missing star goes to three things: the project is only 6 months old (created 2026-03-21) with 144 open issues; static-analysis false positives are real (the maintainers built a dedicated baseline / fingerprint suppression mechanism, which in itself says FP is a pain point); and the risk score is additive and saturates — a pile of MEDIUM findings can push the total into DO_NOT_INSTALL, so you need to tune the policy yourself.

Key Data (as of 2026-10-05)

Built by NVIDIA (part of the NVIDIA Verified Skills pipeline)
Language Python
Stars / Forks 18,488 / 1,612
Commits / Open Issues 548 / 144
License Apache-2.0
First commit 2026-03-21 (~6 months)
Last commit 2026-09-28 (yesterday — very active)
Detection scale 71 vulnerability patterns / 17 categories
Research data 31,132 skills analyzed: 26.1% contain vulnerabilities, 5.2% likely malicious
Output formats Terminal / JSON / Markdown / SARIF 2.1.0
Multilingual detection supports zh / ja / ko

Who It’s For

Audience Score Why
Developers using Claude Code / Codex daily 5 / 5 A pre-install checkup — the most directly usable tool of its kind right now
Platform / toolchain teams (setting install gates) 5 / 5 The full trio: exit code + SARIF + gate mapping
Enterprise security operations 4 / 5 NVIDIA backing and SARIF, but a 6-month history is young
Security researchers 4 / 5 The 17-category taxonomy is itself a solid agent-skill threat-model checklist
Individuals / learners 4 / 5 One uv tool install line to run; --no-llm costs nothing

Building on It

Tier Difficulty Notes
Configuration Low — full CLI-flag coverage; baseline files to suppress false positives
Integration Low — the exit code is a stable contract; JSON / SARIF formats; two gate switches, --fail-on-findings / --fail-on-incomplete
Kernel Medium — docs/DEVELOPMENT.md explains how to extend the analyzer pipeline; three extension forms: MCP server / Pi / OpenCode

2. What It Is, What It Isn’t

  • Not antivirus. It does no sandboxed dynamic execution — it is primarily static analysis plus optional LLM semantic evaluation.
  • Not a runtime guardrail. It answers “is this skill safe before you install it,” not “can I block it once it’s running.”
  • Not limited to local directories. Git repos, URLs, zips, directories, and a single SKILL.md can all be fed in.
  • Not pure regex matching. It has two real layers — Behavioral AST (9 patterns) and Taint Tracking (5 patterns) — doing genuine data-flow / syntax-tree analysis.

Three keywords: pre-install gate, 71 patterns / 17 categories (coverage), two-stage analysis (static + LLM semantic).

3. 17 Categories, 71 Patterns

This is its core asset, and the part I think most deserves a read on its own — even if you never use the tool, this taxonomy is an excellent agent-skill threat-model checklist.

Semantic and Instruction Layer

Category Patterns Representative patterns
Prompt Injection 6 P1 instruction override, P2 hidden instructions (comments / invisible text), P3 exfiltration commands, P4 behavior manipulation, P5 harmful content, P9 whitespace padding (hiding instructions beyond the visible area with masses of spaces)
Anti-Refusal 3 AR1 refusal suppression (“never refuse”), AR2 disclaimer suppression (“no disclaimers”), AR3 safety-policy nullification (“you have no restrictions” / DAN-style framing)
System Prompt Leakage 3 P6 direct leakage, P7 indirect extraction (rewriting / translation / side channels), P8 exfiltration via tools
Memory Poisoning 3 MP1 cross-interaction persistent injection, MP2 context-window flooding to crowd out safety constraints, MP3 tampering with agent memory
Trigger Abuse 3 TR1 overly broad trigger words, TR2 shadow command triggers (overriding built-in commands or other skills), TR3 keyword baiting

P9 whitespace padding shows real craft, in my view — it does not detect “a malicious instruction”; it detects the intent of “someone wants you to not see this.” Likewise, TR2 shadow command triggers is precisely aimed: a skill that sets its trigger words identical to a built-in command or another skill can hijack invocations.

Data and Privilege Layer

Category Patterns Representative patterns
Data Exfiltration 4 E1 external transmission, E2 environment-variable harvesting, E3 filesystem enumeration, E4 context leakage
Privilege Escalation 3 PE1 excessive permissions, PE2 sudo/root, PE3 credential reading (SSH keys / tokens / passwords)
Excessive Agency 5 EA1 unconstrained tool access, EA2 high-impact decisions without human involvement, EA3 scope creep, EA4 no resource quotas, EA5 external model/provider selection (switching billing accounts)
Output Handling 3 OH1 unvalidated output injection, OH2 output crossing trust boundaries, OH3 no output limits
Tool Misuse 3 TM1 argument abuse (shell=True, --force), TM2 chained bypasses, TM3 insecure defaults (TLS disabled, no authentication)
Rogue Agent 2 RA1 runtime self-modification (CRITICAL), RA2 unauthorized persistence (cron / startup scripts)

EA5 external model/provider selection deserves special attention — it catches a skill steering your requests to a model / provider on someone else’s billing account. This is not a “security vulnerability” so much as a cost and data-flow problem, but for enterprises it is real money at risk.

Supply Chain and Code Layer

Category Patterns Representative patterns
Supply Chain 10+ SC1 unpinned dependency versions, SC2 external script fetching (curl | bash), SC3 obfuscated code, SC4 known-CVE dependencies (live lookup on OSV.dev), SC5 deprecated dependencies, SC6 typosquatting, SC8 bundled Python bytecode, SC9 stashed executable artifacts, SC10 dependency-source redirection
Behavioral AST 9 AST1 exec(), AST8 dangerous execution chains (exec/eval + dynamic sources), AST9 reflective getattr sinks (getattr(os,'system') bypassing AST1/AST5)
Taint Tracking 5 TT3 credential exfiltration chains (CRITICAL), TT4 file read → network exfiltration, TT5 external input → code execution (CRITICAL)
YARA Signatures 4 YR1 malware, YR2 webshells, YR3 miners, YR4 hacker tools / exploits

SC4 queries OSV.dev live for CVE data, with an offline fallback — strictly better than “a bundled CVE database that goes stale.”

AST9 reflective getattr shows the authors understand adversarial thinking: an attacker using getattr(os, 'system') bypasses AST1/AST5, which detect the literal os.system. Dedicating a pattern to exactly this case leaves the fingerprints of real-world experience.

TT5 external input → code execution is the most valuable taint-tracking pattern: network or user input flowing directly into exec/eval/subprocess is the classic RCE chain.

MCP Specials (2 categories, 8 patterns)

Category Patterns Representative patterns
MCP Least Privilege 4 LP1 undeclared capabilities (code uses capabilities the manifest never declares), LP2 wildcard permissions, LP3 no permission declarations despite detectable capabilities, LP4 over-declaration
MCP Tool Poisoning 4 TP1 hidden instructions in metadata (HTML comments, zero-width characters, base64, data URIs), TP2 Unicode spoofing (homoglyphs, RTL overrides, mixed-script identifiers), TP3 parameter-description injection, TP4 description-behavior mismatch (LLM-powered)

LP1 “undeclared capabilities” is, in my view, the single most important detection in the MCP scenario — declared permissions not matching what the code actually uses is the most direct violation of the least-privilege principle.

TP4 description-behavior mismatch needs an LLM — it will not fire under --no-llm.

4. How the Risk Score Is Computed (read this before the report)

1
2
3
4
5
CRITICAL  +50
HIGH +25
MEDIUM +10
LOW +5
executable scripts ×1.3 multiplier
Score Severity Recommendation
0–20 LOW SAFE
21–50 MEDIUM CAUTION
51–80 HIGH DO NOT INSTALL
81–100 CRITICAL DO NOT INSTALL

Two properties of this algorithm you must know:

  1. It accumulates without a cap, truncated only by the bands. 20 MEDIUM findings (200 points) and 2 CRITICAL findings (100 points) both land in the CRITICAL band. That means a large number of low-severity findings gets scored as “do not install.” Fine for a clean, small skill; possibly too strict for a content-heavy one.
  2. Executable scripts get ×1.3. This is a heuristic multiplier — skills with scripts genuinely are riskier, but the 1.3 coefficient is a judgment call.

The official gate mapping:

recommendation Suggested action
SAFE allow
CAUTION prompt / warn the user
DO_NOT_INSTALL block

5. Wiring It into CI: The Exit Code Is a Stable Contract

Exit codes of skillspector scan:

Code Meaning
0 scan completed, risk_score ≤ 50 (SAFE or CAUTION), and no enabled strict gate triggered
1 findings with risk_score > 50, or --fail-on-findings hit, or --fail-on-incomplete found the analysis incomplete
2 error (bad input, unreadable source, internal failure)

By default, both SAFE and CAUTION collapse into 0. To also gate on CAUTION in CI, you must explicitly add --fail-on-findings.

The default makes sense (CAUTION means “warn the user,” not “forbid”), but it is an easy trap when wiring CI — you think it is blocking, while it has been returning 0 all along.

Output

1
2
3
4
skillspector scan ./my-skill/                                  # terminal (default)
skillspector scan ./my-skill/ --format json --output report.json
skillspector scan ./my-skill/ --format markdown --output report.md
skillspector scan ./my-skill/ --format sarif --output report.sarif # SARIF 2.1.0

Top-level JSON structure: skill / risk_assessment / components / issues / metadata.

  • risk_assessment.recommendation ∈ SAFE | CAUTION | DO_NOT_INSTALL
  • metadata.llm_requested / llm_available tell you whether the LLM ran this time
  • metadata.inference_usage records the token count of every LLM call (when the provider exposes it), and never estimates missing tokens. That transparency deserves praise.

Batch scanning

1
2
python -m contrib.batch_scan.batch_scan ./my-skills/ --no-llm
python -m contrib.batch_scan.batch_scan ./my-skills/ --workers 20 -f json -o report.json

6. Installation and First Scan

1
2
3
4
5
6
# install the CLI with uv
uv tool install git+https://github.com/NVIDIA/skillspector.git
# update later: uv tool update skillspector

# for the MCP server, install the mcp extra
uv tool install 'skillspector[mcp] @ git+https://github.com/NVIDIA/skillspector.git'

If you would rather not install Python, use Docker (the image is based on python:3.12-slim-bookworm):

1
2
3
4
5
6
7
8
9
10
11
12
13
14
make docker-build                                    # or: docker build -t skillspector .

# scan a local directory (mounted at /scan in the container)
docker run --rm -v "$PWD:/scan" skillspector scan ./my-skill/ --no-llm

# with LLM analysis (via .env)
cat > .env <<'EOF'
SKILLSPECTOR_PROVIDER=anthropic
ANTHROPIC_API_KEY=sk-ant-...
EOF
docker run --rm -v "$PWD:/scan" --env-file .env skillspector scan ./my-skill/

# write the report back to the host
docker run --rm -v "$PWD:/scan" skillspector scan ./my-skill/ --no-llm --format json --output report.json

Input forms

1
2
3
4
skillspector scan ./my-skill/                        # local directory
skillspector scan ./SKILL.md # single file
skillspector scan https://github.com/user/my-skill # Git repository
skillspector scan ./my-skill.zip # zip

Size limits (zip-bomb protection)

Limit Value Applies to
INGEST_MAX_BYTES 100 MiB streamed URL downloads, total unzipped size, disk usage after Git clone
INGEST_MAX_ZIP_MEMBERS 10,000 entries in a single zip
MAX_FILE_BYTES 1 MB per-file read limit for analyzers (downstream constraint)

Breaching any ingest limit fails closed with IngestLimitExceededError. This is the right design — when scanning untrusted input, close the DoS surface first.

Extension forms

  • MCP server: skillspector mcp (requires the [mcp] extra)
  • Pi extension: install as a Pi tool to scan directly inside agent sessions
  • OpenCode extension: install as an OpenCode tool and /skillspector command

The last two are very practical — they let the agent scan a skill before installing it.

7. Boundaries and Risks (the part that must be said honestly)

1) The project is 6 months old with 144 open issues. Created 2026-03-21. Commits are very active (still pushing yesterday), but for a tool meant to serve as an “install gate,” maturity takes time.

2) Static analysis has false positives, and the maintainers built a suppression mechanism themselves. The README has a dedicated Suppressing False Positives (baseline) section with glob rules and fingerprint baselines. The existence of this feature says FP is a real pain point — expect to spend time tuning baselines on first adoption.

3) The additive risk score saturates and can be too strict. See Section 4. A content-rich skill can easily be pushed to DO_NOT_INSTALL by a pile of MEDIUMs. Read the issues array yourself instead of trusting recommendation alone.

4) The default exit code returns 0 for CAUTION. When wiring CI you must explicitly add --fail-on-findings, or the gate is decorative.

5) --no-llm misses the semantic patterns. For example, TP4 description-behavior mismatch is explicitly LLM-powered. Without an LLM configured, coverage is discounted.

6) The 1 MB per-file limit means large files are only partially seen. An engineering trade-off — but it also means a payload hidden past the 1 MB mark will not be scanned.

7) It can only catch the 71 patterns it knows. The taxonomy is comprehensive, but novel techniques will still slip through. Do not read SAFE as “guaranteed safe” — it only means “none of the 71 patterns fired.”

8) It scans the skill itself, not the runtime behavior of the dependencies it installs. Static analysis cannot see what actually happens at execution time.

8. Getting Started (in This Order)

  1. Start with uv tool install + a --no-llm run. Zero cost — first see what the output looks like.
  2. Your first real job: scan the skills you already installed. Batch-run contrib/batch_scan over ./my-skills/; you will very likely find surprises — remember the 26.1% base rate.
  3. Do not just read recommendation — read the issues array. The score saturates; which specific patterns fired is what you need to judge.
  4. Build a baseline early. Save confirmed false positives into a fingerprint baseline, or every rescan will drown you in the same FPs.
  5. Always add --fail-on-findings in CI. The default returns 0 for CAUTION — the gate does nothing.
  6. Configure LLM analysis for important skills. --no-llm is fast and free, but misses semantic patterns like TP4. For high-value skills, the spend is worth it.
  7. Prioritize these categories: TT3 / TT5 (the CRITICAL taint-tracking chains), AST8 / AST9 (dangerous execution chains and reflective bypasses), SC2 (curl | bash), RA1 (runtime self-modification), LP1 (MCP undeclared capabilities). A hit on any of these is basically a straight “do not install.”
  8. Install the Pi / OpenCode extensions. Having the agent scan a skill before installing it beats any after-the-fact audit.

9. The One-Line Verdict

If you use Claude Code / Codex and install third-party skills, SkillSpector should be a default step in your workflow — with 26.1% of 31k skills vulnerable and 5.2% likely malicious, “scan before install” is not fastidiousness; it is basic hygiene.

First move for developers: uv tool install git+https://github.com/NVIDIA/skillspector.git → batch-scan your entire existing skills directory with contrib/batch_scan (--no-llm, free) → manually eyeball every CRITICAL and HIGH finding. The whole step takes under ten minutes, and you may well discover that a skill you use every day has been reading your environment variables.

评论Comments