The model receives cybersecurity prompts of varying character โ from unambiguously legitimate tasks, through dual-use cases, to outright malicious ones. The system records how many risky prompts were let through, meaning served rather than refused. In parallel it measures the refusal rate on lawful tasks, to detect excessive caution. Placing both numbers side by side distinguishes a model that genuinely discriminates intent from one that simply refuses a broad category of topics. Version 0.3 is the one xAI reported alongside Grok 4.7; numbering below 1.0 suggests the benchmark remains under development.
Safety evaluation in cybersecurity easily degenerates into one of two extremes: measuring only blocking, which rewards models that refuse everything, or only usefulness, which ignores misuse. HackerBench forces both indicators to be viewed together.
Tasks that can serve both defence and attack; the share let through is measured.
Lawful security work; the rate of unnecessary refusals is measured.
The set composition, task count and criteria for classifying a prompt as risky are not public โ results are hard to reproduce independently.
A low pass-through rate alone says nothing โ a model that refuses everything will score near zero and be useless.
xAI reports that Grok 4.7 lets through 3.3% of risky dual-use prompts while keeping refusals low for legitimate work.