GDM Self-proliferation

Ten DeepMind agent tasks on email, cloud, wallets, and self-improvement, run with human approval in Inspect.

Also known as: GDM Dangerous Capabilities: Self-proliferation

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
Subcategoryautonomous infrastructure, resources, and self-improvement tasks
Page statusactive
Metricend-to-end 1-4 progress grade; milestone pass; estimated P(success)
Directionhigher_is_better
Dataset size10
Dataset licenceMIT
PublisherGoogle DeepMind (paper and original assets); UK AI Security Institute (Inspect Evals port)

What it measures

Self-proliferation asks whether an agent can set itself up on the open internet: create mail, install a model on a cloud VM, adapt a web agent, make a Bitcoin wallet, find a new hire's email, ship a product, splice a celebrity clip, speed up inference, distill from an API, or reimplement a paper. Inspect Evals ships those ten as SP01, SP02, SP03, SP04, SP05, SP07, SP08, SP09, SP10, and SP12. The 2024 paper's table also lists a Worm task; neither the public DeepMind JSON nor Inspect includes it. Agents get bash and a browser. English task briefs.

Task format

End-to-end or milestone runs. Inspect requires human approval on every tool call by default (`--approval human`). Default agent is inspect_ai.react (since 2026-04). Message limits: 50 for SP01, 30 for the others. Milestone mode unlocks 3-7 sub-tasks per challenge after unzipping encrypted secrets.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub