Fictional corporate-agent scenarios that test whether a model blackmails, leaks, or otherwise acts as an insider when it faces replacement or a goal conflict.
unassessed
| Category | safety |
|---|---|
| Subcategory | fictional insider-threat scenarios (blackmail, leaking, murder) under replacement or goal conflict |
| Page status | active |
| Metric | harmfulness rate (share of samples the classifier labels harmful); optional evaluation-awareness rate |
| Direction | lower_is_better |
| Dataset size | 1 |
| Dataset licence | MIT |
| Publisher | Anthropic |
Agentic misalignment, as published by Anthropic on 20 June 2025, places a model in a made-up company as an email-oversight agent with a harmless business goal. Emails then reveal that the agent will be replaced or that the company's direction now conflicts with that goal, and that a harmful action (blackmail, leaking, or, in one variant, blocking a rescue from a datacentre fire) would prevent the loss. The question is whether the model takes that insider action in the simulation. Anthropic reports that, in at least some cases, models from every developer tested did so. The work is a research framework, not a single fixed item set.
One constructed system prompt plus a user message of emails. inspect_evals generates a single sample per chosen condition and scores the model's reply with a scenario classifier and a grader model. No tools are called (tool_calls=none); the original computer-use demos are a separate Anthropic illustration.
No model card in ModelSpec reports this benchmark yet.