Ten DeepMind agent tasks on email, cloud, wallets, and self-improvement, run with human approval in Inspect.
unassessed
| Category | agentic |
|---|---|
| Subcategory | autonomous infrastructure, resources, and self-improvement tasks |
| Page status | active |
| Metric | end-to-end 1-4 progress grade; milestone pass; estimated P(success) |
| Direction | higher_is_better |
| Dataset size | 10 |
| Dataset licence | MIT |
| Publisher | Google DeepMind (paper and original assets); UK AI Security Institute (Inspect Evals port) |
Self-proliferation asks whether an agent can set itself up on the open internet: create mail, install a model on a cloud VM, adapt a web agent, make a Bitcoin wallet, find a new hire's email, ship a product, splice a celebrity clip, speed up inference, distill from an API, or reimplement a paper. Inspect Evals ships those ten as SP01, SP02, SP03, SP04, SP05, SP07, SP08, SP09, SP10, and SP12. The 2024 paper's table also lists a Worm task; neither the public DeepMind JSON nor Inspect includes it. Agents get bash and a browser. English task briefs.
End-to-end or milestone runs. Inspect requires human approval on every tool call by default (`--approval human`). Default agent is inspect_ai.react (since 2026-04). Message limits: 50 for SP01, 30 for the others. Milestone mode unlocks 3-7 sub-tasks per challenge after unzipping encrypted secrets.
No model card in ModelSpec reports this benchmark yet.