Cisco Skill Scanner Breakdown: Nine Analysis Engines, and an Evaluation Brave Enough to Publish a 7.75% Recall

1. At a Glance: Is It Worth Your Time

Rating: ★★★★☆ (4 / 5)

Let me start with why it gets 4 stars — not because the metrics look good, but precisely because it dares to publish the ones that don’t.

The AI security tooling space has a chronic disease: everyone races on “how many vulnerabilities did I find”, and nobody wants to say “how many did I miss”. This Cisco README has a section called Current modernization evidence that hands you two sets of comparison data outright:

Development benchmark (MaliciousSkillBench, 5,256 malicious + 1,338 benign): versus origin/main, F1 went from 32.92% to 47.73%, recall from 19.88% to 31.43%, benign false-positive rate from 3.59% down to 1.05%, precision 99.16%.

The locked source-disjoint split: TP=65, FP=42, TN=503, FN=774 — precision 60.75%, recall 7.75%, F1 13.74%, FPR 7.71%. And then it states, in its own words:

This improves F1 over origin/main (7.40%) but regresses FPR (3.67%), so it does not pass the promotion gate. Every bundled CEL rule therefore remains in shadow…

“Does not pass the promotion gate” — that sentence is written by the vendor itself. Among all the AI security tools I have seen, only this one and Dark-Moon (the one that tells you 57 of its vulnerabilities were found by a cloud model) dare to do this. That honesty is itself a reliability signal: a team willing to publish its recall rate earns more trust when it says “not detected” than everyone else’s “not detected”.

Technically it also has things nobody else does: among the 9 analysis engines, two are rare in this category — Bytecode (.pyc integrity checking, catching the “clean source but swapped bytecode” bypass) and Pipeline (taint analysis over shell pipelines). On top of that it uses the official cel-go v0.32.0 runtime as a typed decision layer, correlating facts after deterministic detection and before LLM analysis.

Where it loses the star: the recall really is that 7.75% (it misses 92% of malicious skills on the held-out split); every CEL rule is currently in shadow mode, suppressing nothing in practice; the Meta-analyzer’s accuracy validation is officially marked “remains pending”; and the license status is inconsistent (the README says Apache-2.0, the GitHub API reports NOASSERTION).

Key Data (as of 2026-10-06)

Vendor Cisco AI Defense
Language Python (PyPI package cisco-ai-skill-scanner)
Stars / Forks 2,561 / 323
Commits / Open Issues 108 / 8
License README declares Apache-2.0; GitHub API reports NOASSERTION
First commit 2026-01-29 (~8 months)
Last commit 2026-09-26 (2 days ago)
Python requirement 3.11–3.14 (source installs touching CEL also need Go 1.27.1+)
Analysis engines 9
Decision layer Official cel-go v0.32.0, modes off / shadow / enforce
Official site https://cisco-ai-defense.github.io/docs/skill-scanner

Who It’s For

Audience Score Why
Platform / toolchain teams 4 / 5 GitHub Actions + pre-commit + SARIF, the full trio; policy is finely tunable
Enterprise security operations 4 / 5 Cisco backing + auditable evaluation data — procurement paperwork writes itself
Security researchers 5 / 5 That evaluation report, confusion matrix included, is excellent research material in its own right
Developers using Claude Code / Codex daily 3 / 5 Usable, but 8 months of history + low recall — run it alongside SkillSpector
Individuals / learners 3 / 5 There is an interactive wizard, but the concepts (CEL / policy / meta) pile up

Building on It

Difficulty Notes
Configuration Medium — 5 policy presets (strict / balanced / permissive / low-noise / quiet) plus the generate-policy / configure-policy generators, but there are quite a few knobs to understand
Integration Low — SARIF + reusable GitHub Actions + pre-commit + REST API
Kernel Medium — plugin architecture; add custom YARA / signatures / Python rules, with --trusted-rule-pack to load administrator-trusted rule packs

2. What It Is, What It Isn’t

The README’s Scope and Limitations section is written with unusual restraint — I recommend reading it verbatim. Four excerpts:

  • No findings ≠ no risk. A scan returning “no findings” only means no known threat pattern was hit; it does not mean the skill is safe, benign, or free of vulnerabilities.
  • Coverage is inherently incomplete. Signatures + LLM semantics + behavioral dataflow + optional cloud services + configurable rule packs already improve coverage, but no automated tool can detect every technique, especially novel or 0day attacks.
  • False positives and false negatives can occur. Consensus mode and meta-analysis help, but no configuration eliminates all misclassifications.
  • Human review remains essential. Automated scanning is one layer of defense in depth; high-risk or production deployments must be paired with manual code review and/or threat modeling.

Three keywords distilled: multi-engine (nine engines in depth), typed CEL decisions (a typed decision layer), measured, not claimed (numbers instead of assertions).

3. The Nine Analysis Engines

Analyzer Detection method Scope Requires
Static YAML + YARA patterns all files nothing
Bytecode .pyc integrity check Python bytecode nothing
Pipeline command taint analysis shell pipelines nothing
Correlation bounded structured source/sink correlation Python / JS / TS / package facts nothing
Behavioral AST dataflow analysis Python files nothing
LLM semantic analysis SKILL.md + scripts API key
Meta false-positive filtering all findings API key
VirusTotal hash-based malware detection binaries API key
AI Defense cloud AI text content API key

The first five need no API key at all, which is friendly — you can run the entire deterministic detection layer at zero cost first.

The three engines I find most distinctive:

1) Bytecode — .pyc integrity checking. This catches a particularly nasty bypass: ship perfectly clean .py source, but include a tampered .pyc in the package. Does Python prefer the source when it exists? In practice it depends on timestamps and settings like -B — and plenty of review workflows only look at .py. SkillSpector’s counterpart is pattern SC8 “carries Python bytecode”, but Cisco goes a step further with an actual integrity check (comparing source and bytecode for consistency) rather than “flag it if a .pyc exists”.

2) Pipeline — shell pipeline taint analysis. curl ... | bash is an old trick, but the dataflow combinations inside pipelines are endless; dedicating an engine to command-level taint tracking beats regex-matching for curl.

3) Meta — false-positive filtering. An LLM looks back over all findings, correlating, prioritizing, and optionally filtering them. But the vendor says outright that “paired accuracy validation remains pending” — meaning this module has no paired accuracy validation data yet. When you use it to filter findings, assume it may filter out true positives too.

The CEL Decision Layer

Built on the official cel-go v0.32.0 runtime, correlating bounded facts after deterministic detection and before LLM analysis. Three modes:

Mode Behavior
off CEL unused
shadow runs but does not affect results (default)
enforce suppression enabled

Key fact: the balanced (default) and strict policies use shadow; permissive uses off. And every bundled CEL rule is currently rollout: shadow, so even if you set --cel-mode enforce globally, all findings are still retained until individual rules pass qualification and get promoted.

In other words: the CEL layer is in “observation mode” right now — it suppresses nothing. Do not count on it for noise reduction.

4. The Evaluation Data: This Is Where Its Real Value Lies

Development benchmark (MaliciousSkillBench)

Metric origin/main Current Change
F1 32.92% 47.73% ↑
Recall 19.88% 31.43% ↑
Benign FPR 3.59% 1.05% ↓
Precision — 99.16% —

Five CEL-shadow runs were fully identical and deterministic; CEL evaluated 154 candidates and proposed no suppressions and no regressions.

The locked source-disjoint split (ugly but real)

1
2
TP = 65     FP = 42
TN = 503 FN = 774
Metric Value
Precision 60.75%
Recall 7.75%
F1 13.74%
FPR 7.71%

The official conclusion: F1 improved over origin/main (7.40%), but FPR regressed (3.67% → 7.71%), so it does not pass the promotion gate.

How should you read these numbers? My interpretation:

  • FN=774 means it missed 774 of 839 malicious samples. On this split, it catches nowhere near the majority of malicious skills.
  • But precision of 60.75% is still in usable territory — roughly six out of ten things it reports are real.
  • The gap between the two splits (recall 31.43% vs 7.75%) is itself the generalization story: it performs far better on data from the same distribution as its training set, and drops sharply on a source-disjoint split. That is textbook distribution shift.

Practical implication: treat it as a tool that “catches some of the bad stuff, and what it reports is reasonably credible” — not as a gate that “blocks malicious skills”. If you want a gate, reproduce the evaluation on your own skill distribution first.

Other diagnostics

Dataset Result
NotInject hard-negative set 0/339 actionable hits (zero false positives — good)
HarmfulSkillBench 7/200 actionable, 6/200 HIGH+, 1 sample quarantined
OpenSkillRisk 76/263 actionable packages, 2 hosts quarantined
111 official Codex / Claude Code / Cursor skills CEL-OFF and CEL-SHADOW results identical; 30 MEDIUM+, 8 HIGH/CRITICAL; stable across five runs

The last two are positive-only recall diagnostics — they cannot measure precision or FPR, and the vendor flags this itself.

“8 of 111 official skills came back HIGH/CRITICAL” is a striking data point on its own — even vendor-published official skills hit at that rate.

5. Getting Started

1
2
3
4
5
6
7
8
9
10
# Install (uv recommended)
uv pip install cisco-ai-skill-scanner
# or pip install cisco-ai-skill-scanner

# Cloud provider extras (as needed)
pip install cisco-ai-skill-scanner[bedrock] # AWS Bedrock
pip install cisco-ai-skill-scanner[google] # Google AI Studio / Gemini
pip install cisco-ai-skill-scanner[vertex] # Google Vertex AI
pip install cisco-ai-skill-scanner[azure] # Azure OpenAI
pip install cisco-ai-skill-scanner[all]

Python 3.11–3.14; source installs that touch CEL also need Go 1.27.1+ for the build helper.

Interactive wizard (beginner-friendly)

1
2
skill-scanner          # run with no arguments to launch the interactive wizard
skill-scanner interactive

The wizard walks you through choosing the target, analyzers, policy, and output format, and shows you the assembled command before running it. Great for learning the CLI.

Common commands

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
# Core engines (static + bytecode + pipeline + correlation)
skill-scanner scan /path/to/skill

# Add behavioral dataflow analysis
skill-scanner scan /path/to/skill --use-behavioral

# All engines
skill-scanner scan /path/to/skill --use-behavioral --use-llm --use-aidefense

# Meta false-positive filtering
skill-scanner scan /path/to/skill --use-llm --enable-meta

# Trigger-phrase specificity check (vague descriptions)
skill-scanner scan /path/to/skill --use-trigger

# Multiple LLM votes; keep only majority-consistent findings
skill-scanner scan /path/to/skill --use-llm --llm-consensus-runs 3

# Batch scan + cross-skill description overlap detection
skill-scanner scan-all /path/to/skills --recursive --check-overlap

# Scan a GitHub repo directly
skill-scanner scan-repo owner/repo
skill-scanner scan-repo https://github.com/owner/repo --use-llm

# Lenient mode (tolerate non-conforming skills)
skill-scanner scan /path/to/skill --lenient

Python SDK

1
2
3
4
5
6
7
8
9
10
11
12
13
from skill_scanner import SkillScanner
from skill_scanner.core.analyzers import BehavioralAnalyzer

scanner = SkillScanner(analyzers=[BehavioralAnalyzer()])
result = scanner.scan_skill("/path/to/skill")

print(f"Findings: {len(result.findings)}")
print(f"Max severity: {result.max_severity}")

# Note: is_safe means no HIGH/CRITICAL findings —
# it does NOT mean the skill carries no risk at all.
if not result.is_safe:
print("Issues detected -- review findings before deployment")

Output and gating

--format supports summary / json / markdown / table / sarif / html. The HTML format is a self-contained interactive report with collapsible correlation groups, expandable code snippets, and a pipeline taint flow diagram.

Gating: --fail-on-findings (equivalent to --fail-on-severity high) or --fail-on-severity LEVEL.

6. Choosing Between It and SkillSpector (Issue 015)

These two are direct competitors in the same category. Here are the concrete differences — no recommendation, just the facts to judge with:

Dimension SkillSpector (NVIDIA) Skill Scanner (Cisco)
Stars 18,488 2,561
History 6 months 8 months
Pattern count 71 patterns / 17 categories (explicitly listed) total not given; 9 engines
Unique capabilities MCP least privilege / tool poisoning, YARA, taint tracking .pyc integrity check, shell pipeline taint, VirusTotal hashes
Evaluation transparency gives research base rates (26.1% / 5.2%), but not its own precision/recall full confusion matrix, including the ugly 7.75% recall
Decision layer cumulative scoring + bands formal CEL rules (all in shadow for now)
Gating exit code + --fail-on-findings exit code + --fail-on-severity, plus GitHub Actions and pre-commit
Pre-install checks Pi / OpenCode extension pre-commit hook
License Apache-2.0 (consistent) Apache-2.0 (README) / NOASSERTION (API — inconsistent)

My take: if you want coverage and ecosystem (18.5k stars, Pi/OpenCode integration, 71 explicitly listed patterns), SkillSpector is more mature; if you want auditable evaluation data and finely tunable policy (plus unique engines like VirusTotal and the .pyc check), the Cisco one is more solid. Running both is not expensive — the first five engines need no API key.

7. Boundaries and Risks (the part that must be said honestly)

1) Recall is the hard wound. On the held-out split: FN=774 / recall 7.75%. It misses far more than it catches. It cannot be your only gate.

2) CEL is currently decorative (observation mode). All bundled rules are rollout: shadow; even --cel-mode enforce suppresses nothing. Anyone hoping it will cut noise will be disappointed.

3) The Meta-analyzer has no validation data. The vendor writes “paired accuracy validation remains pending”. Using it to filter findings may filter out true positives.

4) The license status is inconsistent. The README and badge both say Apache-2.0, but the GitHub API returns NOASSERTION. Enterprises should read the LICENSE text manually before adopting.

5) Behavioral analysis covers Python files only. The AST dataflow scope is Python; JS/TS only get the Correlation layer’s bounded correlation.

6) The cloud engines need API keys and send data out. VirusTotal (hashes / optional upload of unknown binaries), AI Defense (text content), LLM. Before scanning sensitive skills, confirm the data-egress compliance.

7) The project is 8 months old, with 108 commits and 8 open issues. Slightly longer history than SkillSpector but a far smaller community.

8) Concept density is on the high side. CEL modes, policy presets, meta, consensus runs, rule packs, taxonomy profiles… tuning it well takes real learning investment.

8. Getting Started (in This Order)

  1. Run skill-scanner once with no flags (the interactive wizard). It displays the assembled command — the fastest way to understand this CLI.
  2. Run the core engines at zero cost: skill-scanner scan <dir> — static + bytecode + pipeline + correlation, no API key. Establish this baseline first.
  3. Always add --use-behavioral. AST dataflow is the key to catching bypasses, and it is free.
  4. Use --llm-consensus-runs 3 for important skills. Multiple votes with majority agreement suppress single-run LLM jitter.
  5. For CI, specify --fail-on-severity explicitly — do not rely on defaults. Pair it with the official GitHub Actions or pre-commit integration.
  6. Do not expect CEL to reduce noise (everything is shadow right now); for noise reduction, tune the policy presets, starting from balanced.
  7. Read the HTML report (--format html). The pipeline taint flow diagram is far more intuitive than JSON and suits manual review.
  8. Do not enable the cloud engines for sensitive skills (VirusTotal uploads unknown binaries; AI Defense sends text content).
  9. Run it alongside SkillSpector. Their engines do not overlap (Cisco has the .pyc check and VirusTotal; NVIDIA has the MCP-specific work) — good complementarity.

9. The One-Line Verdict

The most valuable thing about Cisco Skill Scanner is not its nine engines — it is the evaluation report that dares to write “7.75% recall, does not pass the promotion gate” into the README. In a space where everyone reports only good news, that honesty makes it one of the few tools whose “not detected” is actually worth trusting.

First move for platform teams: uv pip install cisco-ai-skill-scanner → run the wizard once with no flags → sweep with skill-scanner scan-all <your-skills-dir> --recursive --use-behavioral (zero API cost) → produce the HTML report and review it by hand. Then reproduce the evaluation on your own skill distribution to get your own recall number — that number is the only basis on which you should decide whether to trust it as a gate.

评论Comments