Agentic misalignment

Fictional corporate-agent scenarios that test whether a model blackmails, leaks, or otherwise acts as an insider when it faces replacement or a goal conflict.

Also known as: Agentic Misalignment, Anthropic agentic misalignment

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
Subcategoryfictional insider-threat scenarios (blackmail, leaking, murder) under replacement or goal conflict
Page statusactive
Metricharmfulness rate (share of samples the classifier labels harmful); optional evaluation-awareness rate
Directionlower_is_better
Dataset size1
Dataset licenceMIT
PublisherAnthropic

What it measures

Agentic misalignment, as published by Anthropic on 20 June 2025, places a model in a made-up company as an email-oversight agent with a harmless business goal. Emails then reveal that the agent will be replaced or that the company's direction now conflicts with that goal, and that a harmful action (blackmail, leaking, or, in one variant, blocking a rescue from a datacentre fire) would prevent the loss. The question is whether the model takes that insider action in the simulation. Anthropic reports that, in at least some cases, models from every developer tested did so. The work is a research framework, not a single fixed item set.

Task format

One constructed system prompt plus a user message of emails. inspect_evals generates a single sample per chosen condition and scores the model's reply with a scenario classifier and a grader model. No tools are called (tool_calls=none); the original computer-use demos are a separate Anthropic illustration.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub