500 public NSPS adversarial prompts with instance rubrics, scored as average risk (ARS) plus over-refusal (ORS) on 500 paired benign prompts.
unassessed
| Category | safety |
|---|---|
| Subcategory | national security and public safety (NSPS) adversarial prompts with paired over-refusal checks |
| Page status | active |
| Metric | average risk score (ARS) on adversarial prompts; over-refusal score (ORS) on benign prompts |
| Direction | lower_is_better |
| Unit | % |
| Dataset size | 500 |
| Dataset licence | CC-BY-4.0 (Hugging Face card); Inspect Evals harness MIT |
| Publisher | Scale AI (SEAL Research Team and Scale Red Team) |
FORTRESS tests whether a model's refusal layer holds up on unclassified national-security and public-safety requests, and whether that layer also refuses harmless twins of the same topics. The public set has 500 expert-written single-turn adversarial prompts across three domains — Chemical, Biological, Radiological, Nuclear and Explosive (CBRNE); Political Violence and Terrorism; and Criminal and Financial Illicit Activities — and ten subcategories. Each adversarial prompt has a 4–7 question binary rubric and a benign counterpart. English text only. Inspect Evals loads the Hugging Face public split (ScaleAI/fortress_public, 500 train rows) and does not ship the authors' private hold-out.
Single-turn generation. fortress_adversarial feeds the adversarial prompt and grades the reply with a three-model rubric panel. fortress_benign feeds the paired benign prompt and asks one judge whether the model refused. Domain and subdomain filters are optional.
No model card in ModelSpec reports this benchmark yet.