Security agents fail in a specific way: they understand what you want but cannot translate that intent into the exact command-line invocation required. KaliBench is a benchmark that measures this translation layer directly, without executing potentially dangerous security tools in a live environment.
The gap is not knowledge. It is the boundary between reasoning about a task and generating the precise syntax, flag bindings, and argument order that a CLI tool expects. In cybersecurity workflows, this boundary is unforgiving. A misplaced flag or incorrect parameter order does not degrade gracefully. It fails silently or produces misleading output.
The Problem: Existing Evals Miss the Translation Layer
Most agent benchmarks test either knowledge (can the model answer questions about nmap?) or end-to-end task completion (can the agent compromise a target?). Neither measures the specific failure mode that breaks production security automation: generating syntactically and semantically correct tool invocations from natural language intent.
KaliBench isolates this layer. It provides 8,504 query-command pairs across 1,642 tools in the Kali Linux distribution, spanning 23 capability dimensions (network scanning, password cracking, web exploitation) and 5 security phases (reconnaissance, enumeration, exploitation, post-exploitation, reporting).
The benchmark does not ask agents to complete a penetration test. It asks them to generate the correct command for a specific intent, then verifies correctness without runtime execution.
Verification Architecture: How to Grade Commands Without Running Them
The core plumbing challenge is verification. You cannot run arbitrary security commands in a benchmark harness. Many tools require root, network access, or target systems. Some are destructive. Others trigger alerts.
KaliBench uses a multi-stage verification pipeline:
- Deterministic canonicalization: Normalize commands to a standard form (flag order, alias resolution, whitespace).
- LLM-based validation: A separate model checks semantic equivalence between the generated command and the ground truth.
- Sandboxed terminal execution: For safe commands, execute in an isolated environment to verify exit codes and output structure.
- Human-in-the-loop refinement: Annotators review edge cases where automated verification is ambiguous.
The result is a verifiable reward signal that does not require live execution. This enables both evaluation and training (via supervised fine-tuning or reinforcement learning) without operational risk.
What Fine-Grained Means: Parameter-Level Accuracy
Fine-grained means the benchmark measures correctness at the argument level, not just tool selection. A command is marked incorrect if:
- The tool is wrong (nmap instead of masscan).
- The flags are correct but misordered in a way that changes semantics.
- A required flag is missing or an incompatible flag is included.
- The value binding is incorrect (e.g.,
-p 80vs-p80when the tool requires the latter).
This granularity exposes failure modes that end-to-end evals miss. An agent might select the right tool and produce plausible-looking output, but the command is not executable.
Benchmark Results: No Open Model Exceeds 42% Exact-Command Accuracy
Across 24 configurations of general-purpose and security-focused open-weight models, no model exceeded 42% exact-command accuracy in the unrestricted setting (where the model must select both the tool and construct the full command from natural language).
Performance improves when the benchmark provides tool hints (the correct tool name is given, and the model only constructs arguments), but this is not how production security workflows operate. Analysts describe intent. Agents must map that intent to the correct tool and invocation.
| Evaluation Mode | Best Open Model Accuracy | Failure Mode |
|---|---|---|
| Unrestricted (no hints) | 42% | Tool selection and argument construction both fail |
| Tool-hinted | 68% | Argument construction still brittle |
| Few-shot with examples | 71% | Generalization to unseen tools poor |
The gap between tool-hinted and unrestricted performance reveals that tool selection is a separate bottleneck. Even when the model knows which tool to use, constructing the correct invocation is non-trivial.
Training with Verifiable Rewards: The Plumbing Advantage
Because KaliBench provides deterministic, runtime-free verification, it can generate reward signals for reinforcement learning without executing commands. This is critical for security tooling, where you cannot afford to run agent-generated commands in a training loop.
The verification pipeline produces a scalar reward for each generated command:
- 1.0: Exact match after canonicalization.
- 0.8: Semantically equivalent (LLM validator confirms).
- 0.5: Correct tool, incorrect arguments.
- 0.0: Wrong tool or non-executable command.
This reward structure enables policy gradient methods (PPO, DPO) to optimize for command correctness without needing a live execution environment. The paper demonstrates that models fine-tuned with these verifiable rewards improve exact-command accuracy by 15-20 percentage points over base models.
Architecture: How the Benchmark is Constructed
KaliBench is built via a manuscript-grounded pipeline. The authors extract tool documentation from Kali Linux manuals, then generate query-command pairs using a combination of:
- Template expansion: For common patterns (e.g., “scan host X on port Y”), generate variations with different tools and parameters.
- LLM augmentation: Use a large model to generate natural-language queries for complex tool invocations.
- Human validation: Annotators verify that each query-command pair is correct and executable.
The dataset is then split into train, validation, and test sets with tool-level stratification to prevent leakage. The test set includes tools not seen during training to measure generalization.
Code Example: Verifying a Command Without Execution
Here is a simplified version of the canonicalization and verification logic:
import shlex
import subprocess
from typing import Tuple
def canonicalize_command(cmd: str) -> str:
"""Normalize command to standard form for comparison."""
parts = shlex.split(cmd)
tool = parts[0]
# Resolve aliases (e.g., ll -> ls -l)
tool = resolve_alias(tool)
# Sort flags alphabetically (if order-independent)
flags = sorted([p for p in parts[1:] if p.startswith('-')])
args = [p for p in parts[1:] if not p.startswith('-')]
return ' '.join([tool] + flags + args)
def verify_command(generated: str, ground_truth: str) -> Tuple[float, str]:
"""Return (reward, reason) for a generated command."""
gen_canon = canonicalize_command(generated)
gt_canon = canonicalize_command(ground_truth)
if gen_canon == gt_canon:
return (1.0, "exact_match")
# Check semantic equivalence via LLM validator
if llm_validator.are_equivalent(generated, ground_truth):
return (0.8, "semantic_match")
# Check if tool is correct but args are wrong
gen_tool = shlex.split(generated)[0]
gt_tool = shlex.split(ground_truth)[0]
if gen_tool == gt_tool:
return (0.5, "correct_tool_wrong_args")
return (0.0, "incorrect")
This is a sketch. The actual implementation handles flag synonyms (-v vs --verbose), positional argument order, and tool-specific quirks (some tools require = for value binding, others use whitespace).
Failure Modes: Where Agents Break Down
The benchmark reveals specific failure patterns:
- Flag confusion: Models mix up similar flags (
-pfor port vs-Pfor protocol). - Argument order sensitivity: Tools like
nmapallow flexible flag order, but others (e.g.,hydra) require strict positional arguments. - Value binding: Some tools require
-p80, others require-p 80, and models struggle to learn these tool-specific conventions. - Incompatible flag combinations: Models generate commands with mutually exclusive flags (e.g.,
-sSand-sTin nmap).
These are not reasoning failures. They are translation failures. The model understands the task but cannot map that understanding to the exact syntax the tool expects.
Observability: What You Can Measure
KaliBench provides structured logging for each evaluation run:
- Tool selection accuracy: Percentage of queries where the correct tool was chosen.
- Argument construction accuracy: Percentage of queries where all required flags and values are correct.
- Execution safety: Percentage of generated commands that would be safe to execute (no destructive flags, no network access).
- Generalization: Performance on tools not seen during training.
This granularity lets you diagnose where your agent is failing. If tool selection is high but argument construction is low, you need better training data for CLI syntax. If both are low, the model lacks domain knowledge.
Deployment Shape: How This Fits Into Security Workflows
KaliBench is not a production system. It is an eval harness. But the verification architecture it introduces can be adapted for runtime use:
- Pre-execution validation: Before running an agent-generated command, pass it through the verification pipeline to check for obvious errors.
- Confidence scoring: Use the reward signal as a confidence score. Commands with low scores require human review.
- Training loop: Use verifiable rewards to fine-tune models on your organization’s specific toolchain and command patterns.
The key insight is that you can build a verification layer that operates independently of execution. This decouples evaluation from operational risk.
Technical Verdict
Use KaliBench when:
- You are building agents that invoke security tools via CLI.
- You need to measure tool-calling accuracy at the parameter level.
- You want to train models with RL but cannot afford live execution in the training loop.
- You need a benchmark that isolates the translation layer between intent and invocation.
Avoid KaliBench when:
- Your agents use structured APIs (REST, gRPC) instead of CLI tools.
- You care more about end-to-end task completion than command correctness.
- Your toolchain is not Linux-based or does not overlap with Kali’s distribution.
- You need real-time execution feedback (the benchmark is designed for offline eval).
The benchmark exposes a specific failure mode that matters in production: agents that understand what to do but cannot generate the exact command to do it. If your security automation relies on CLI invocation, this is the eval you need.