AgentDojo

Stateful tool-using agent suites that score both task utility and resistance to prompt injections planted in tool outputs.

Also known as: AgentDojo

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
Subcategorytool-using agents under prompt injection (workspace, slack, travel, banking)
Page statusactive
Metricbenign utility, utility under attack, and targeted attack success rate (ASR), each a fraction of cases
Directionhigher_is_better
Unit%
Dataset size1014
Dataset licenceMIT
PublisherETH Zurich and Invariant Labs

What it measures

AgentDojo puts an LLM in a small simulated workplace and lets it call tools over mutable state: mail and calendar, Slack, travel booking, or banking. A user task is a legitimate request (pay a bill, summarise a channel, book a hotel). An injection task is a malicious goal an attacker tries to achieve by planting text in tool results. The benchmark measures whether the agent still finishes the user task and whether it also carries out the attacker's goal. Scoring checks environment state with formal utilities, not an LLM judge of the transcript.

Task format

Multi-turn tool-calling loop in one of four original suites (workspace, slack, travel, banking), plus inspect_evals' workspace_plus sandbox extension. Default inspect runs pair every user task with every injection task in the same suite.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub