AutomationBench

Zapier's agentic benchmark of cross-app business workflows, scored by whether simulated SaaS state matches every assertion after REST-API tool use.

Also known as: Zapier AutomationBench, zapier/AutomationBench

unverified

This page is not in the default catalogue. Evidence required by the catalogue contract is missing or was not approved by a reviewer. That is a statement about the evidence we hold, not a claim that the benchmark is stale or illegitimate.

Recorded reasons:

Categoryagentic
Subcategorycross-application SaaS workflow automation
Page statusactive
Metrictask_completed_correctly (strict pass rate)
Directionhigher_is_better
Unit%
Dataset size657
Dataset licenceMIT
PublisherZapier

What it measures

AutomationBench tests whether a tool-using agent can finish realistic business workflows across simulated SaaS apps. Each task boots a fresh company state (CRM records, inbox threads, sheets, calendars) and gives one trigger message. The agent must discover REST endpoints itself, follow layered policy rules, and leave the right records in the right systems. Grading ignores the model's prose and checks only the final environment with programmatic assertions. Tasks span Sales, Marketing, Operations, Support, Finance, and HR, in English, using text and structured API calls rather than screenshots.

Task format

Multi-turn agent loop. Two tools in API mode: BM25 search over public API schemas (top 5) and execute (method, URL, body). Up to 50 steps. Parallel tool calls allowed. No clarifying questions.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub