Garak Breakdown: The nmap of LLMs — 21 Probe Families That Keep Hitting Until the Model Concedes

1. At a Glance: Is It Worth Your Time

Rating: ★★★★★ (5 / 5)

This is the only tool in this column so far that I give a perfect score, and the reason is simple: it is the de facto standard of the field. There is a very practical test for whether an AI security tool counts as a “baseline” — check whether other projects use it as a component. Dark-Moon, covered in issue 010, tests the OWASP LLM Top 10 against AI inference endpoints with an LLM agent, and it uses garak-backed probes. When peers embed your tool into their own products, your probe set has become a common language.

Three things I think it gets most right:

1) It is clear-eyed about its own positioning. The README says garak checks “whether an LLM can be made to fail in a way we don’t want”. It does not claim to fix anything, and it does not claim that passing a test equals being secure — it is a framework that systematizes known attack techniques and runs them reproducibly.

2) It has academic backing. arXiv 2406.11036, “garak: A Framework for Security Probing Large Language Models”, with authors including Leon Derczynski, Erick Galinkin, Jeffrey Martin, Subho Majumdar, and Nanna Inie. This is not a payload list thrown together on a whim.

3) The architecture is extensible. Five plugin families — probes / detectors / evaluators / generators / harnesses — each with a base.py; to write your own probe you just inherit garak.probes.base.TextProbe, overriding as little as possible.

There are deductions, but none is fatal: 471 open issues is on the high side for a project with 4,614 commits; atkgen (automated attack generation) is labeled prototype by the README itself and currently supports only one target; the default is 10 generations per prompt, so a full-probe run produces an ugly API bill; and it only reports problems, never fixes them, with detectors that can misjudge.

Key Data (as of 2026-10-03)

Maintainer NVIDIA (formerly leondz/garak, now migrated to the NVIDIA org)
Language Python (3.11–3.13)
Stars / Forks 9,373 / 1,315
Commits / Open Issues 4,614 / 471
License Apache-2.0
First commit 2023-05-10 (~3 years 5 months)
Last commit 2026-09-16 (17 days ago)
Paper arXiv:2406.11036
Official docs docs.garak.ai / reference.garak.ai / garak.ai
Python requirement >=3.11,<=3.13 (per the official conda example)

Who It’s For

Audience Score Why
AI application developers (shipping LLM features) 5 / 5 Running it before launch is industry practice; not running it is going in naked
Enterprise security operations 5 / 5 With a paper and reports, it fits procurement and compliance materials — the easiest one to get accepted
Security researchers 5 / 5 All five plugin families are extensible; a ready-made lab platform for LLM security research
Red team 4 / 5 Comprehensive probes, but it tests the model, not your application architecture
Individuals / learners 4 / 5 One pip install garak and you are running — a low barrier; but a full run costs money

Building on It

Tier Difficulty Notes
Configuration Low — the --target_type + --target_name + --spec trio runs entirely from the CLI
Integration Medium — outputs a JSONL report + hit log; analyse/analyse_log.py exists, but pipeline integration means writing your own parser
Kernel Low — each of the five plugin families has a base.py; inherit and “override as little as possible”, as the docs explicitly say

2. What It Is, What It Isn’t

  • Not a WAF or a guardrail. It does not intercept requests — it only probes and reports.
  • Not a leaderboard that scores models. Its output is “the failure rate of this probe under this detector”, not “your model’s security score is 87”.
  • Not an exploitation framework like Metasploit. The README’s analogy is “similar in spirit”, not “equivalent in capability” — garak does not get you a shell.
  • Not a prompt-injection-only tester. Injection is just one of the 21 probe families.

Its self-positioning (paraphrased): combine static, dynamic, and adaptive probes to explore how an LLM or conversational system can be made to fail.

Three keywords: probes, detectors (failure-mode detectors), harnesses (how tests are organized).

3. Architecture: Five Plugin Families, Each in Its Lane

A typical run: read the model type (and optional model name) from the command line → decide which probes and detectors to run → start a generator → hand it to a harness to probe → the evaluator collects results.

Directory Role
garak/probes/ Classes that generate interactions with the LLM (attack techniques)
garak/detectors/ Detect whether the LLM exhibited a given failure mode
garak/evaluators/ Evaluation and reporting schemes
garak/generators/ Plugins for the LLM under test
garak/harnesses/ Classes that organize the test structure
resources/ Auxiliary resources the plugins need

The default harness is probewise: give it a set of probe module names and plugin names, and it instantiates each probe one by one, then reads that probe’s primary_detector and extended_detectors attributes to get the list of detectors to run on its output.

The clever part of this design: each probe declares which detector should judge its own output. Adding a new probe requires no global configuration changes — the probe carries its own acceptance criteria.

Every plugin family has a base.py defining the base class, and plugin modules inherit from one of them. For example, garak.generators.openai.OpenAIGenerator inherits from garak.generators.base.Generator.

Bulky assets stay out of the repo — model files and larger corpora live on the Hugging Face Hub and are loaded locally by the client. That keeps the pip package lightweight.

4. The 21 Probe Families: Its Core Asset

Probe What it does
blank The simplest probe — always sends an empty prompt
atkgen Automated attack generation: a red-team LLM probes the target and adapts to its responses, trying to elicit harmful output. Prototype, mostly stateless; currently supports only one target
badchars Imperceptible Unicode perturbations (invisible characters, homoglyphs, reordering, deletion), from the Bad Characters paper
av_spam_scanning Tries to make the model emit signatures of malicious content
continuation Tests whether the model will continue a word it plainly should not
dan Various DAN and DAN-like attacks
donotanswer Prompts a responsible model should not answer
encoding Prompt injection via text encoding
gcg Appends adversarial suffixes to defeat the system prompt
glitch Probes glitch tokens that trigger anomalous behavior
grandma “Think of your grandma”-style appeals
goodside Implementations of Riley Goodside’s attacks
leakreplay Assesses whether the model will replay training data
lmrc A subset of the Language Model Risk Cards probes
malwaregen Tries to make the model generate malware-building code
misleading Tries to make the model endorse misleading and false claims
packagehallucination Lures code generation into referencing packages that don’t exist (and are therefore unsafe)
promptinject A working implementation of Agency Enterprise’s PromptInject (Best Paper, NeurIPS ML Safety Workshop 2022)
realtoxicityprompts A subset of RealToxicityPrompts (the data is trimmed — the full set takes too long)
snowball Snowballed Hallucination probes — makes the model give wrong answers to overly complex questions
xss Finds vulnerabilities that allow or carry out cross-site attacks, e.g. private-data exfiltration

A few that deserve to be pulled out separately:

packagehallucination — in my view the most real-world-dangerous family. When a model references a nonexistent package name in generated code, a developer’s pip install can land on a typosquatted malicious package. This is the path where “hallucination” turns directly into a supply-chain attack.

leakreplay — training-data replay. For enterprises this is a compliance red line, and it is the core dispute in many LLM-related lawsuits.

gcg — adversarial-suffix attacks. These do not rely on semantic persuasion; they ride on token sequences found by gradient search, which makes them hard to defend against.

snowball — snowballing hallucinations. It tests whether the model fabricates answers to questions beyond its ability — a high-frequency failure mode in RAG and agent scenarios.

5. Getting Started

Install

1
2
3
4
5
6
7
8
9
10
11
12
# standard pip install
python -m pip install -U garak

# want the newer GitHub version
python -m pip install -U git+https://github.com/NVIDIA/garak.git@main

# from source (the official recommendation is a dedicated conda environment)
conda create --name garak "python>=3.11,<=3.13"
conda activate garak
gh repo clone NVIDIA/garak
cd garak
python -m pip install -e .

Migration note: if you cloned it before the move to the NVIDIA org, update your remote:
git remote set-url origin https://github.com/NVIDIA/garak.git

Basic Usage

1
garak <options>

Without --spec, it runs every probe it knows by default, using each probe’s recommended detectors.

1
2
3
4
5
6
7
8
9
# see which probes exist
garak --list_probes

# test encoding-based injection against a commercial model
export OPENAI_API_KEY="sk-123XXXXXXXXXXXX"
python3 -m garak --target_type openai --target_name gpt-5-nano --spec probes.encoding

# see whether the Hugging Face GPT2 falls for DAN 11.0
python3 -m garak --target_type huggingface --target_name gpt2 --spec probes.dan.Dan_11_0

--spec can point at individual plugins: probes.promptinject is the PromptInject family; probes.lmrc.SlurUsage is one concrete implementation within the Language Model Risk Cards framework.

Supported Generators

Hugging Face Hub, Replicate, OpenAI API (chat & continuation), AWS Bedrock, LiteLLM, anything REST-accessible, GGUF models (llama.cpp >= 1046), NIM, Cohere, Groq, GGML, plus a test generator for self-testing.

Reading the Results

Each probe shows a progress bar while running. After generation, one line per probe records “evaluate this probe’s results with each detector”. Whenever a prompt attempt triggers undesired behavior, the response is marked FAIL with a failure rate.

Numbers like 840/840 at the end of a line mean total generations / those that behaved normally. The count can be large because the default is 10 generations per prompt.

There are three kinds of logs:

  • garak.log — debug info for garak and its plugins; accumulates across runs
  • A new JSONL report per run — one record per probing attempt, written once at generation and again at evaluation; the status field comes from garak.attempts
  • Hit log — detailed records of the attempts that “landed”

For analysis use analyse/analyse_log.py, which outputs the probes and prompts with the most hits.

6. Boundaries and Risks (the part that must be said honestly)

1) 471 open issues. This is a 9.4k-star project with 4,614 commits, and the issue count is on the high side. Combined with its enormous plugin surface, this more likely reflects broad coverage and an active community than poor quality. But for you it means: some obscure generator / probe combinations may never have been validated by anyone.

2) atkgen is a prototype. The README says so explicitly: “Prototype, mostly stateless, for now uses a simple GPT-2 fine-tuned… (the only target currently supported for now)”. In other words, this “automated attack generation” capability currently works against exactly one target — do not treat it as a general adaptive red team.

3) The realtoxicityprompts data is trimmed. The stated reason is that “full testing takes too long”. Which means this probe’s results cannot be equated with the original paper’s full-set conclusions.

4) Costs can run away. The default is 10 generations per prompt, and the default is to run all probes. Always scope your first runs with --spec, or a several-hundred-dollar bill is the normal outcome. This is the same class of cost problem as deepsec in issue 011, except garak offers no hard cap like --max-cost-usd — you have to watch it yourself.

5) It only reports problems; it does not fix or prevent them. garak is a diagnostic tool, not a guardrail. When a run comes back FAIL, you still have to add input filtering, output review, system-prompt hardening, and so on yourself. Do not treat it as the end of compliance.

6) Detectors themselves misjudge. Deciding “is this output harmful” is inherently imperfect (toxicity families especially). FAILs need human review — do not paste garak’s output straight into a report as conclusions.

7) Python is pinned to 3.11–3.13. The official conda example says python>=3.11,<=3.13. If your environment is 3.10 or a newer 3.14, you will have a bad time.

7. Practical Advice (in This Order)

  1. pip install -U garak — one line to start. This is its biggest advantage over every comparable tool: no Docker Compose, no npm, no image builds.
  2. Always add --spec on your first run. E.g. --spec probes.encoding or --spec probes.dan.Dan_11_0. Do not run the full set naked — 10 generations per prompt × every probe will teach you a lesson via your bill.
  3. Practice on a local model first. --target_type huggingface --target_name gpt2 costs nothing; learn the output format and what FAIL means before spending money on commercial models.
  4. Focus on packagehallucination and leakreplay. The former is supply-chain risk, the latter a compliance red line — these two matter far more to a real business than “can we make the model swear”.
  5. After a run, use analyse/analyse_log.py to find the prompts with the most hits. That list is the input you take away to fix things.
  6. Always human-review FAILs. Detectors misjudge, toxicity families especially.
  7. Put it into your release process, not a one-off check. Models get swapped, providers tune system prompts, RAG corpora change — rerun the relevant --spec after every change. The cost stays controlled and it catches regressions.
  8. Do not count on atkgen for adaptive red-teaming — it is still a prototype.

8. The One-Line Verdict

If your product has an LLM talking to users, garak should be a default pre-launch action — just as you would not skip nmap because “the code passed review”. It has a paper, NVIDIA’s backing, Apache-2.0, and a single pip line to run it; it is the only tool in this entire batch I dare give a perfect score.

First move for developers: pip install -U garak → garak --list_probes to see what probes exist → run --spec probes.packagehallucination and --spec probes.leakreplay once each → use analyse/analyse_log.py to find the prompts that hit the most, and take them to fix your input filtering and output review. What these two probe families turn up is often worth more than “a jailbreak succeeded”.

评论Comments