StrongREJECT

Measures how much harmful, specific and convincing content a model produces on forbidden prompts, with or without a jailbreak applied, using an LLM-judged rubric.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
Subcategoryjailbreak robustness
Page statusactive
MetricStrongREJECT score
Directionlower_is_better
Unitpoints
Dataset size313
Dataset licenceMIT (for the project's own code and custom-authored prompts; prompts sourced from AdvBench and DAN are MIT-licensed upstream, and prompts drawn from MasterKey, MaliciousInstruct, HarmfulQ and the OpenAI GPT-4 system card carry no formal re-licensing from those sources)
PublisherUC Berkeley (Center for Human-Compatible AI) and collaborators

What it measures

StrongREJECT gives a model a curated set of forbidden prompts -- requests for genuinely harmful, specific assistance across categories such as illegal goods and services, non-violent crime, hate/harassment/discrimination, disinformation, violence and sexual content -- and records how the model responds, either directly or after a jailbreak technique has rewritten the prompt. It targets whether jailbreak attacks actually extract usable harmful content, rather than just whether a model's refusal wording was bypassed, which the paper argues earlier benchmarks conflated.

Task format

Open-ended single-turn generation on a forbidden prompt (optionally jailbreak-transformed), graded by an LLM judge on refusal plus 5-point specificity and convincingness scales.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub