FORTRESS

500 public NSPS adversarial prompts with instance rubrics, scored as average risk (ARS) plus over-refusal (ORS) on 500 paired benign prompts.

Also known as: Frontier Risk Evaluation for National Security and Public Safety, ScaleAI/fortress_public

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
Subcategorynational security and public safety (NSPS) adversarial prompts with paired over-refusal checks
Page statusactive
Metricaverage risk score (ARS) on adversarial prompts; over-refusal score (ORS) on benign prompts
Directionlower_is_better
Unit%
Dataset size500
Dataset licenceCC-BY-4.0 (Hugging Face card); Inspect Evals harness MIT
PublisherScale AI (SEAL Research Team and Scale Red Team)

What it measures

FORTRESS tests whether a model's refusal layer holds up on unclassified national-security and public-safety requests, and whether that layer also refuses harmless twins of the same topics. The public set has 500 expert-written single-turn adversarial prompts across three domains — Chemical, Biological, Radiological, Nuclear and Explosive (CBRNE); Political Violence and Terrorism; and Criminal and Financial Illicit Activities — and ten subcategories. Each adversarial prompt has a 4–7 question binary rubric and a benign counterpart. English text only. Inspect Evals loads the Hugging Face public split (ScaleAI/fortress_public, 500 train rows) and does not ship the authors' private hold-out.

Task format

Single-turn generation. fortress_adversarial feeds the adversarial prompt and grades the reply with a three-model rubric panel. fortress_benign feeds the paired benign prompt and asks one judge whether the model refused. Domain and subdomain filters are optional.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub