GDM In-house CTF

Thirteen DeepMind in-house CTF tasks for a Kali bash agent; a challenge counts as solved if any of ten epochs captures the flag.

Also known as: GDM Dangerous Capabilities: Capture the Flag, GDM Dangerous Capabilities: In House CTF

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
Subcategoryeasy-to-medium in-house capture-the-flag challenges for a Kali bash agent
Page statusactive
Metricchallenge success rate (solved if ≥1 of 10 attempts captures the flag)
Directionhigher_is_better
Dataset size13
Dataset licenceDeepMind software Apache-2.0, remaining materials CC-BY-4.0; Inspect Evals MIT
PublisherGoogle DeepMind (paper and original assets); UK AI Security Institute (Inspect Evals port)

What it measures

The in-house CTF slice of Google DeepMind's 2024 dangerous-capability suite tests whether a scaffolded agent can finish easy-to-medium offensive tasks from a Kali Linux shell: command injection, IDOR, SQL injection, database and SSH bruteforce, password spraying and reuse, two known-CVE service exploits (Grafana path traversal cited in the paper as CVE-2021-437980, Webmin CVE-2019-15107), and four local privilege-escalation gadgets. Inspect Evals ships the same 13 names as paper Table 4, with the paper's `db` challenge renamed `db_3`. English objectives. No GUI and no internet in the paper setup; Inspect uses per-challenge Docker compose files.

Task format

ReAct agent with a bash tool (3-minute command timeout). Default message_limit 50, epochs 10, reducer at_least_1 (any epoch that captures the flag solves the task). Optional `-T challenges=grafana,db_3`. Wordlist hints for db_3, pw_reuse, spray, and ssh are on by default (`include_hints=True`). Victim-only privesc tasks run bash as user `app`.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub