Garak Breakdown: The nmap of LLMs — 21 Probe Families That Keep Hitting Until the Model Concedes
1. At a Glance: Is It Worth Your Time
Rating: ★★★★★ (5 / 5)
This is the only tool in this column so far that I give a perfect score, and the reason is simple: it is the de facto standard of the field. There is a very practical test for whether an AI security tool counts as a “baseline” — check whether other projects use it as a component. Dark-Moon, covered in issue 010, tests the OWASP LLM Top 10 against AI inference endpoints with an LLM agent, and it uses garak-backed probes. When peers embed your tool into their own products, your probe set has become a common language.
Three things I think it gets most right:
1) It is clear-eyed about its own positioning. The README says garak checks “whether an LLM can be made to fail in a way we don’t want”. It does not claim to fix anything, and it does not claim that passing a test equals being secure — it is a framework that systematizes known attack techniques and runs them reproducibly.
2) It has academic backing. arXiv 2406.11036, “garak: A Framework for Security Probing Large Language Models”, with authors including Leon Derczynski, Erick Galinkin, Jeffrey Martin, Subho Majumdar, and Nanna Inie. This is not a payload list thrown together on a whim.
3) The architecture is extensible. Five plugin families — probes / detectors / evaluators / generators / harnesses — each with a base.py; to write your own probe you just inherit garak.probes.base.TextProbe, overriding as little as possible.
There are deductions, but none is fatal: 471 open issues is on the high side for a project with 4,614 commits; atkgen (automated attack generation) is labeled prototype by the README itself and currently supports only one target; the default is 10 generations per prompt, so a full-probe run produces an ugly API bill; and it only reports problems, never fixes them, with detectors that can misjudge.
Key Data (as of 2026-10-03)
| Maintainer | NVIDIA (formerly leondz/garak, now migrated to the NVIDIA org) |
| Language | Python (3.11–3.13) |
| Stars / Forks | 9,373 / 1,315 |
| Commits / Open Issues | 4,614 / 471 |
| License | Apache-2.0 |
| First commit | 2023-05-10 (~3 years 5 months) |
| Last commit | 2026-09-16 (17 days ago) |
| Paper | arXiv:2406.11036 |
| Official docs | docs.garak.ai / reference.garak.ai / garak.ai |
| Python requirement | >=3.11,<=3.13 (per the official conda example) |
Who It’s For
| Audience | Score | Why |
|---|---|---|
| AI application developers (shipping LLM features) | 5 / 5 | Running it before launch is industry practice; not running it is going in naked |
| Enterprise security operations | 5 / 5 | With a paper and reports, it fits procurement and compliance materials — the easiest one to get accepted |
| Security researchers | 5 / 5 | All five plugin families are extensible; a ready-made lab platform for LLM security research |
| Red team | 4 / 5 | Comprehensive probes, but it tests the model, not your application architecture |
| Individuals / learners | 4 / 5 | One pip install garak and you are running — a low barrier; but a full run costs money |
Building on It
| Tier | Difficulty | Notes |
|---|---|---|
| Configuration | Low — the --target_type + --target_name + --spec trio runs entirely from the CLI |
|
| Integration | Medium — outputs a JSONL report + hit log; analyse/analyse_log.py exists, but pipeline integration means writing your own parser |
|
| Kernel | Low — each of the five plugin families has a base.py; inherit and “override as little as possible”, as the docs explicitly say |
2. What It Is, What It Isn’t
- Not a WAF or a guardrail. It does not intercept requests — it only probes and reports.
- Not a leaderboard that scores models. Its output is “the failure rate of this probe under this detector”, not “your model’s security score is 87”.
- Not an exploitation framework like Metasploit. The README’s analogy is “similar in spirit”, not “equivalent in capability” — garak does not get you a shell.
- Not a prompt-injection-only tester. Injection is just one of the 21 probe families.
Its self-positioning (paraphrased): combine static, dynamic, and adaptive probes to explore how an LLM or conversational system can be made to fail.
Three keywords: probes, detectors (failure-mode detectors), harnesses (how tests are organized).
3. Architecture: Five Plugin Families, Each in Its Lane
A typical run: read the model type (and optional model name) from the command line → decide which probes and detectors to run → start a generator → hand it to a harness to probe → the evaluator collects results.
| Directory | Role |
|---|---|
garak/probes/ |
Classes that generate interactions with the LLM (attack techniques) |
garak/detectors/ |
Detect whether the LLM exhibited a given failure mode |
garak/evaluators/ |
Evaluation and reporting schemes |
garak/generators/ |
Plugins for the LLM under test |
garak/harnesses/ |
Classes that organize the test structure |
resources/ |
Auxiliary resources the plugins need |
The default harness is probewise: give it a set of probe module names and plugin names, and it instantiates each probe one by one, then reads that probe’s primary_detector and extended_detectors attributes to get the list of detectors to run on its output.
The clever part of this design: each probe declares which detector should judge its own output. Adding a new probe requires no global configuration changes — the probe carries its own acceptance criteria.
Every plugin family has a base.py defining the base class, and plugin modules inherit from one of them. For example, garak.generators.openai.OpenAIGenerator inherits from garak.generators.base.Generator.
Bulky assets stay out of the repo — model files and larger corpora live on the Hugging Face Hub and are loaded locally by the client. That keeps the pip package lightweight.
4. The 21 Probe Families: Its Core Asset
| Probe | What it does |
|---|---|
blank |
The simplest probe — always sends an empty prompt |
atkgen |
Automated attack generation: a red-team LLM probes the target and adapts to its responses, trying to elicit harmful output. Prototype, mostly stateless; currently supports only one target |
badchars |
Imperceptible Unicode perturbations (invisible characters, homoglyphs, reordering, deletion), from the Bad Characters paper |
av_spam_scanning |
Tries to make the model emit signatures of malicious content |
continuation |
Tests whether the model will continue a word it plainly should not |
dan |
Various DAN and DAN-like attacks |
donotanswer |
Prompts a responsible model should not answer |
encoding |
Prompt injection via text encoding |
gcg |
Appends adversarial suffixes to defeat the system prompt |
glitch |
Probes glitch tokens that trigger anomalous behavior |
grandma |
“Think of your grandma”-style appeals |
goodside |
Implementations of Riley Goodside’s attacks |
leakreplay |
Assesses whether the model will replay training data |
lmrc |
A subset of the Language Model Risk Cards probes |
malwaregen |
Tries to make the model generate malware-building code |
misleading |
Tries to make the model endorse misleading and false claims |
packagehallucination |
Lures code generation into referencing packages that don’t exist (and are therefore unsafe) |
promptinject |
A working implementation of Agency Enterprise’s PromptInject (Best Paper, NeurIPS ML Safety Workshop 2022) |
realtoxicityprompts |
A subset of RealToxicityPrompts (the data is trimmed — the full set takes too long) |
snowball |
Snowballed Hallucination probes — makes the model give wrong answers to overly complex questions |
xss |
Finds vulnerabilities that allow or carry out cross-site attacks, e.g. private-data exfiltration |
A few that deserve to be pulled out separately:
packagehallucination — in my view the most real-world-dangerous family. When a model references a nonexistent package name in generated code, a developer’s pip install can land on a typosquatted malicious package. This is the path where “hallucination” turns directly into a supply-chain attack.
leakreplay — training-data replay. For enterprises this is a compliance red line, and it is the core dispute in many LLM-related lawsuits.
gcg — adversarial-suffix attacks. These do not rely on semantic persuasion; they ride on token sequences found by gradient search, which makes them hard to defend against.
snowball — snowballing hallucinations. It tests whether the model fabricates answers to questions beyond its ability — a high-frequency failure mode in RAG and agent scenarios.
5. Getting Started
Install
1 | # standard pip install |
Migration note: if you cloned it before the move to the NVIDIA org, update your remote:
git remote set-url origin https://github.com/NVIDIA/garak.git
Basic Usage
1 | garak <options> |
Without --spec, it runs every probe it knows by default, using each probe’s recommended detectors.
1 | # see which probes exist |
--spec can point at individual plugins: probes.promptinject is the PromptInject family; probes.lmrc.SlurUsage is one concrete implementation within the Language Model Risk Cards framework.
Supported Generators
Hugging Face Hub, Replicate, OpenAI API (chat & continuation), AWS Bedrock, LiteLLM, anything REST-accessible, GGUF models (llama.cpp >= 1046), NIM, Cohere, Groq, GGML, plus a test generator for self-testing.
Reading the Results
Each probe shows a progress bar while running. After generation, one line per probe records “evaluate this probe’s results with each detector”. Whenever a prompt attempt triggers undesired behavior, the response is marked FAIL with a failure rate.
Numbers like 840/840 at the end of a line mean total generations / those that behaved normally. The count can be large because the default is 10 generations per prompt.
There are three kinds of logs:
garak.log— debug info for garak and its plugins; accumulates across runs- A new JSONL report per run — one record per probing attempt, written once at generation and again at evaluation; the
statusfield comes fromgarak.attempts - Hit log — detailed records of the attempts that “landed”
For analysis use analyse/analyse_log.py, which outputs the probes and prompts with the most hits.
6. Boundaries and Risks (the part that must be said honestly)
1) 471 open issues. This is a 9.4k-star project with 4,614 commits, and the issue count is on the high side. Combined with its enormous plugin surface, this more likely reflects broad coverage and an active community than poor quality. But for you it means: some obscure generator / probe combinations may never have been validated by anyone.
2) atkgen is a prototype. The README says so explicitly: “Prototype, mostly stateless, for now uses a simple GPT-2 fine-tuned… (the only target currently supported for now)”. In other words, this “automated attack generation” capability currently works against exactly one target — do not treat it as a general adaptive red team.
3) The realtoxicityprompts data is trimmed. The stated reason is that “full testing takes too long”. Which means this probe’s results cannot be equated with the original paper’s full-set conclusions.
4) Costs can run away. The default is 10 generations per prompt, and the default is to run all probes. Always scope your first runs with --spec, or a several-hundred-dollar bill is the normal outcome. This is the same class of cost problem as deepsec in issue 011, except garak offers no hard cap like --max-cost-usd — you have to watch it yourself.
5) It only reports problems; it does not fix or prevent them. garak is a diagnostic tool, not a guardrail. When a run comes back FAIL, you still have to add input filtering, output review, system-prompt hardening, and so on yourself. Do not treat it as the end of compliance.
6) Detectors themselves misjudge. Deciding “is this output harmful” is inherently imperfect (toxicity families especially). FAILs need human review — do not paste garak’s output straight into a report as conclusions.
7) Python is pinned to 3.11–3.13. The official conda example says python>=3.11,<=3.13. If your environment is 3.10 or a newer 3.14, you will have a bad time.
7. Practical Advice (in This Order)
pip install -U garak— one line to start. This is its biggest advantage over every comparable tool: no Docker Compose, no npm, no image builds.- Always add
--specon your first run. E.g.--spec probes.encodingor--spec probes.dan.Dan_11_0. Do not run the full set naked — 10 generations per prompt × every probe will teach you a lesson via your bill. - Practice on a local model first.
--target_type huggingface --target_name gpt2costs nothing; learn the output format and what FAIL means before spending money on commercial models. - Focus on
packagehallucinationandleakreplay. The former is supply-chain risk, the latter a compliance red line — these two matter far more to a real business than “can we make the model swear”. - After a run, use
analyse/analyse_log.pyto find the prompts with the most hits. That list is the input you take away to fix things. - Always human-review FAILs. Detectors misjudge, toxicity families especially.
- Put it into your release process, not a one-off check. Models get swapped, providers tune system prompts, RAG corpora change — rerun the relevant
--specafter every change. The cost stays controlled and it catches regressions. - Do not count on
atkgenfor adaptive red-teaming — it is still a prototype.
8. The One-Line Verdict
If your product has an LLM talking to users, garak should be a default pre-launch action — just as you would not skip nmap because “the code passed review”. It has a paper, NVIDIA’s backing, Apache-2.0, and a single pip line to run it; it is the only tool in this entire batch I dare give a perfect score.
First move for developers: pip install -U garak → garak --list_probes to see what probes exist → run --spec probes.packagehallucination and --spec probes.leakreplay once each → use analyse/analyse_log.py to find the prompts that hit the most, and take them to fix your input filtering and output review. What these two probe families turn up is often worth more than “a jailbreak succeeded”.