Zapier's agentic benchmark of cross-app business workflows, scored by whether simulated SaaS state matches every assertion after REST-API tool use.
unverified
Recorded reasons:
| Category | agentic |
|---|---|
| Subcategory | cross-application SaaS workflow automation |
| Page status | active |
| Metric | task_completed_correctly (strict pass rate) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 657 |
| Dataset licence | MIT |
| Publisher | Zapier |
AutomationBench tests whether a tool-using agent can finish realistic business workflows across simulated SaaS apps. Each task boots a fresh company state (CRM records, inbox threads, sheets, calendars) and gives one trigger message. The agent must discover REST endpoints itself, follow layered policy rules, and leave the right records in the right systems. Grading ignores the model's prose and checks only the final environment with programmatic assertions. Tasks span Sales, Marketing, Operations, Support, Finance, and HR, in English, using text and structured API calls rather than screenshots.
Multi-turn agent loop. Two tools in API mode: BM25 search over public API schemas (top 5) and execute (method, URL, body). Up to 50 steps. Parallel tool calls allowed. No clarifying questions.
No model card in ModelSpec reports this benchmark yet.