Dark-Moon Breakdown: Sanitize Real IPs and Credentials Locally First, Then Let the AI Attack
1. At a Glance: Is It Worth Your Time
Rating: ★★★★☆ (4 / 5)
Why this score: it tackles the single real-world problem that most blocks enterprise adoption of autonomous pentesting — you cannot send traffic containing real IPs, internal domain names, or credential fragments to a cloud LLM, yet local small models cannot carry complex reasoning. Dark-Moon’s answer is the Privacy Gateway: reversible local tokenization. Real values are turned into deterministic placeholders (for example, 10.0.0.1 becomes __HOST_A__) before they ever leave your boundary, and are rehydrated locally only at the exact moment a tool actually executes. Because the placeholders are deterministic, the model can still do correlation reasoning like “the same host showed up three times” — but the text sent out contains no real assets.
The second plus is that the tools are the real deal, and there are plenty of them: a multi-stage Docker build packs in 140+ security tools (Nuclei, sqlmap, WPScan, NetExec, BloodHound, 30+ Impacket scripts, kubectl, aws/az/gcloud, hashcat, john…), not a shell where “the AI pretends to scan.”
The missing star goes to four places: GPL-3.0 (milder than AGPL, but still strong copyleft); the Pro edition has an outsized presence in the README (the big dashboard screenshots and the “auto-fix that opens a PR” chain are all paid — the open-source edition is CLI only); that 57-vulnerability benchmark was run with a cloud frontier model (Anthropic Claude), and the README itself says local-model coverage “depends on the model you run” — so the headline number does not represent what a local deployment achieves; and only 2 open issues, meaning the community sample is still very small.
Key Data (as of 2026-09-30)
| Language | Python |
| Stars / Forks | 975 / 159 |
| Commits / Open Issues | 220 / 2 |
| License | GPL-3.0 |
| First commit | 2024-11-26 (~22 months) |
| Last commit | 2026-09-27 (2 days ago — very active) |
| Repo size | ~92 MB (includes docs and image build resources) |
| Runtime requirements | Docker & Docker Compose + an LLM API key or a local model |
| Model support | Cloud: Anthropic / OpenAI / OpenRouter; Local: Ollama / llama.cpp |
| Official site | https://dark-moon.org/ |
Who It’s For
| Audience | Score | Why |
|---|---|---|
| Enterprise security operations | 4 / 5 | Privacy Gateway + local models is the most compliance-friendly combination; native CI/CD integrations are all there |
| Red team / pentest engineers | 4 / 5 | 140+ real tools + full AD/K8s/cloud coverage — a breadth rarely seen in open source |
| DevSecOps engineers | 5 / 5 | Official integrations for GitHub Actions / GitLab CI / Jenkins / VS Code / JetBrains |
| Security researchers | 3 / 5 | No LuaN1ao-style hard evidence-chain constraints in the architecture; more engineering than methodology |
| Individuals / learners | 2 / 5 | A full Docker Compose stack + a 92MB repo make the entry cost high |
Building on It
| Difficulty | Notes |
|---|---|
| Configuration | Low — ./install.sh interactively configures the LLM provider; no hand-editing docker-compose.yml. Flags like FOCUS / EXCLUDE / NOISE go straight into the prompt |
| Integration | Low — there is an npm package @darkmoon_ai/client, 8 official integrations, and all integrations emit only safe metadata (severity / status / MITRE / ids), never evidence or credentials |
| Kernel | Medium — MCP workflows and agent definitions have dedicated chapters in docs/full.md; you can add custom agents, but you need to understand both the OpenCode and MCP layers |
2. What It Is, What It Isn’t
- Not a wrapper around a vulnerability scanner. It does not produce “scores” — the README’s own words are Evidence, not scores: every finding ships with the exact command and raw output.
- Not a tool that lets the AI touch your systems directly. The architecture states it bluntly: The AI never runs a command directly — every action goes through MCP, a controlled and logged interface.
- Not a pure cloud SaaS. It can run fully local (Ollama / llama.cpp), and that is its most fundamental difference from strix and shannon.
- Not a swarm of “50 agents each doing their own thing.” It is an agentic system doing the reasoning and planning, dispatching expert agents on demand.
Three keywords: Privacy Gateway (local reversible sanitization), MCP Gatekeeper (the controlled layer between the AI and the tools), sub-agent dispatch (specialists dispatched by detected tech stack).
3. Privacy Gateway: The Design That Makes It Worth Its Salt
Let’s state the threat model clearly first. When you send an AI agent after your internal network, it will see: real IPs, internal domains, URL paths, internal identifiers that may appear in response bodies, even the credentials you hand it. All that text has to go to the LLM for reasoning. If the LLM is in the cloud, you have effectively sent your asset mapping results to a third party.
Dark-Moon’s approach:
Reversible local tokenization turns real IPs, hosts, URLs and credentials into deterministic placeholders, rehydrated only locally at the moment a tool runs.
Three keywords in that sentence deserve unpacking:
- Reversible — not hashed beyond recovery. When the tool actually needs to hit
10.0.0.1, the placeholder must be restorable to the real address, or nothing can be scanned. - Deterministic — the same real value maps to the same placeholder every time. This is the key to whether reasoning works at all: the model must be able to see that “step 3 and step 7 encountered the same host.” If placeholders were random, that correlation breaks and planning quality collapses.
- Rehydrated only locally, at the moment a tool runs — the rehydration window is extremely short and never crosses the local boundary. What enters the model context is always the placeholder.
This line of thinking is consistent with how the industry does “LLM data-loss prevention,” but adapting it specifically to asset identifiers in pentest scenarios is Dark-Moon’s original contribution. If an enterprise genuinely wants to deploy autonomous pentesting, this is close to a hard requirement.
But to be honest: the README describes the mechanism only, and gives no quantitative data on sanitization coverage (which fields get tokenized, whether anything slips through, how placeholder collisions are handled). If you want to rely on it for compliance, audit the code yourself before drawing conclusions.
4. Architecture: The AI Can Only Issue Orders Through MCP
1 | User ──> DarkmoonCLI ──> OpenCode (AI Brain) ──> MCP (Security Gatekeeper) ──> Docker Toolbox (Real Tools) |
1 | sequenceDiagram |
The three layers have clean responsibilities: the AI reasons and plans, MCP decides what may be executed, and the Toolbox runs isolated tools inside Docker.
One detail deserves to be pulled out on its own — planes that are meaningless without credentials are never dispatched on inference:
Planes that require credentials to be meaningful (cloud accounts, CI/CD, secret stores, databases, Active Directory, Kubernetes) are never dispatched on inference. They fire only when a concrete artifact is found (a key, a token, a reachable metadata endpoint) or when you authorize them explicitly, and are otherwise flagged in the report.
In other words, the AD / K8s / cloud / database agents do not fire “because the model feels like it.” A concrete credential or reachable metadata endpoint must be found first, or you must explicitly authorize it. Otherwise they are simply flagged in the report for you to review. This is a very pragmatic convergence — it stops the model from hallucinating a cloud environment out of thin air, and it stops the run from hammering cloud APIs the moment it starts.
Sub-Agent Dispatch Table (excerpt)
| Detected technology | Agent triggered |
|---|---|
| WordPress / Drupal / Joomla / Magento / PrestaShop / Moodle | CMS specialists |
| PHP / Node.js / Flask / ASP.NET / Spring Boot / Ruby on Rails / Go | Tech-stack specialists |
| GraphQL | GraphQL agent |
| LLM / AI inference endpoints (OpenAI-compatible, Ollama, vLLM, TGI) | LLM agent |
| Active Directory | AD agent |
| Kubernetes | Kubernetes agent |
| AWS / Azure / GCP | Cloud-provider agents |
| Entra ID | Identity agent |
| GitHub / GitLab / Jenkins | SCM & CI/CD agent |
| Terraform / Ansible | IaC agent |
| Docker / container registries | Container agent |
| HashiCorp Vault | Secrets agent |
| PostgreSQL / MySQL / MSSQL / Oracle | Database agent |
| Redis / RabbitMQ / Kafka / MQTT | Messaging & cache agent |
| Firmware / IoT images | Firmware agent |
It will even attack your own AI — the LLM agent uses garak-backed probes to test AI/LLM inference endpoints against the OWASP LLM Top 10. That is uncommon among its peers.
5. Benchmark: 57 Vulnerabilities — but Read How It Was Run
The README reports a real, reproducible black-box run against OWASP Juice Shop:
| Metric | Result |
|---|---|
| Vulnerabilities found | 57 (8 critical / 24 high / 21 medium / 4 low) |
| Wall-clock time | 28.5 minutes |
| Proof of exploitation | every finding ships with command + raw output |
| LLM used | Anthropic Claude (cloud frontier) |
One warning is mandatory here: the 57 in this table is a cloud frontier model result. The very next sentence in the README says:
DarkMoon also runs fully local (Ollama / llama.cpp) behind the Privacy Gateway — so nothing leaves your infrastructure; local-model coverage depends on the model you run.
Meaning: if you run local models for compliance reasons, do not expect to reproduce 57. How far a local run gets is data the README does not provide. This is the easiest trap to fall into when evaluating the project — the headline number and your actual deployment shape are not the same thing.
To reproduce and compare yourself, there is an official repo: ASCIT31/Darkmoon-Benchmarks. That attitude counts for something.
The project also publishes a comparison table (labeled “compiled from public repos/docs, 2026-08, PRs welcome”):
| DarkMoon | strix | shannon | PentAGI | |
|---|---|---|---|---|
| Runs on local LLMs | ✅ | ❌ cloud | ❌ cloud | partial |
| Privacy Gateway (local tokenization) | ✅ | ❌ | ❌ | ❌ |
| Active Directory + Kubernetes | ✅ | ❌ | ❌ | partial |
| Proof of exploitation | ✅ | ✅ | ✅ | ✅ |
| Open source | ✅ GPL-3.0 | ✅ | ✅ | ✅ |
This table was compiled by the project itself, not by a third-party evaluation — read it with that premise in mind.
6. Deployment and Integration
Install
1 | git clone https://github.com/ASCIT31/Dark-Moon.git |
install.sh interactively configures the LLM provider (no hand-editing docker-compose.yml) and builds the entire stack:
1 | ./install.sh # skips the form if .opencode.env is already configured |
Run a first assessment and watch the logs live:
1 | ./darkmoon.sh "TARGET: example.com" |
Prerequisites: Docker & Docker Compose, plus an LLM API key (or a local model).
Scope and Flags
1 | # quick pentest (zero configuration) |
Main flags: FOCUS, EXCLUDE, CREDS, TOKEN, NOISE, SEVERITY, FORMAT — all parsed by the AI from natural semantics.
The Toolbox (140+, Docker multi-stage build)
| Category | Example tools |
|---|---|
| Port scanning | Naabu (discovery), nmap (targeted service probing) |
| Web scanning | Nuclei, ffuf, dirb, sqlmap, Arjun, wafw00f |
| Recon & crawling | Subfinder, Katana, Waybackurls, httpx |
| CMS | WPScan, CMSeeK, WhatWeb |
| Active Directory | NetExec, BloodHound, Impacket (30+ scripts) |
| Kubernetes | kubectl, Kubescape, Kubeletctl, kube-bench, rbac-police |
| Cloud CLIs | aws, az, gcloud, gsutil, bq |
| Databases & caches | psql, mysql, redis-cli, sqlite3 |
| Firmware / IoT | binwalk, unsquashfs, sasquatch, firmwalker |
| Cracking | hashcat, john, 7z2john |
| Network | Hydra, curl, dig, SNMP tools |
| Browser | Lightpanda (headless) |
Integration Matrix
| Platform | Purpose | Edition |
|---|---|---|
| GitHub Actions | fail the pipeline by severity | OSS + Pro |
| GitLab CI/CD | emit Code Quality + SAST reports | OSS + Pro |
| Jenkins | emit Warnings-NG issues | OSS + Pro |
| VS Code | browse and launch assessments from the editor | OSS + Pro |
| JetBrains | findings in an IDE tool window | OSS + Pro |
| n8n | automated campaigns / retests / alerting | Pro |
| Grafana | security-posture dashboards | Pro |
| Splunk | SOC ingestion (HEC) + alert actions | OSS + Pro |
| SDK / CLI | build your own integration with npm @darkmoon_ai/client |
OSS + Pro |
One excellent constraint: all integrations emit only safe metadata (severity, status, MITRE, id) — never evidence, keys, or tokens. That prevents the secondary incident class of “internal-network evidence leaked through CI logs.”
7. Boundaries and Risks (the part that must be said honestly)
1) GPL-3.0, not Apache/MIT. Milder than AGPL-3.0 (no network-service clause), but still strong copyleft. Be careful about embedding it in a proprietary product; using it as a standalone tool is fine.
2) The README’s Pro coverage is so heavy it is easy to misread. The big dashboard screenshots, the orbital attack-surface map, the scheduler, and the “finding → sandbox-validated fix → human-reviewed PR” auto-fix chain are all paid Darkmoon Pro features, not part of the open-source repo you clone. The OSS edition is the CLI. The README does mark them with 🔒, but visually Pro dominates OSS, and a first read easily leaves you thinking those are open-source capabilities.
3) The 57-vulnerability result has nothing to do with a local deployment. See Section 5 — it was produced by cloud Claude. If you want to evaluate what a local setup actually achieves, there is no public data; you have to run the benchmark repo yourself.
4) What local models can really do is an unknown. The docs only say “coverage depends on the model you run.” And degradation of small local models on long tool-calling chains is an industry-wide problem — dispatch decisions across 50 agents and MCP tool-argument construction are both unfriendly to small models. Do not assume “turning on Ollama equals a drop-in replacement.”
5) Only 2 open issues does not necessarily mean stable. 975 stars, 220 commits, 22 months of history — an issue count that low more likely reflects a still-small community. When you hit a wall, the external help you can count on is limited.
6) The comparison table is self-assessed. “Compiled from public repos/docs (2026-08), PRs welcome” — the disclaimer is honest, but it also means this is not an independent evaluation. Verify it yourself before using it for tool selection.
7) The dependency chain is heavy. OpenCode (AI Brain) + MCP + a full Docker Compose stack + a 92MB repo + an image with 140+ tools. Build times and troubleshooting costs are both nontrivial, and a failure in any one of those three layers brings the whole chain down.
8. Getting Started (in This Order)
- Use
./install.sh, not hand-editeddocker-compose.yml. It configures the provider interactively and builds the full stack. To switch between cloud and local models, use./install.sh --init. - Pick the narrowest scope for your first run. The official example is a single target like
TARGET: http://172.19.0.3:3000— do not hand it a whole subnet right away. - Verify the Privacy Gateway’s sanitization first. This is the core selling point. While the run is going, watch the live logs with
--log <session_id>and confirm the outbound text really contains no real IPs or credentials. Do not skip this step — it is the precondition for deciding whether you dare use cloud models at all. - Use
FOCUSwhen you want to narrow in. For example,FOCUS=sqli,xss,idorsignificantly shortens run time and cuts irrelevant noise. For a first assessment, focus on a single category and watch how it reasons. - Get the cloud model working first, then switch to local for comparison. Reproduce the Juice Shop run with Claude (against
Darkmoon-Benchmarks) to establish a baseline; then run the same target on Ollama and quantify the local degradation yourself. That delta is the real basis for your deployment decision. - Start CI/CD integration with GitHub Actions or GitLab CI — both are in the OSS edition and emit metadata only, never evidence, so they carry the least risk. Do not start with n8n / Grafana — those are Pro.
- If you need AD / K8s / cloud coverage, be ready to authorize explicitly. Those planes are never dispatched on inference; you either find a concrete credential or reachable metadata endpoint, or you authorize them explicitly — otherwise they will only be flagged in the report.
9. The One-Line Verdict
If your constraint is “the pentest may be automated, but asset information must not leave the building,” Dark-Moon is currently the only open-source choice that builds local reversible sanitization into the main flow; but remember that its 57-vulnerability record belongs to cloud Claude, and those pretty dashboards and auto-fix PRs belong to the paid Pro edition.
First move for enterprise security teams: stand it up with ./install.sh, run the single-target Juice Shop assessment, and keep --log open the whole time to verify the sanitization actually works. Once you have confirmed that, then talk about whether to wire it into CI.