Artificial Analysis's independent run of Zapier's private AutomationBench split, scoring objective completion with zero credit after any guardrail break.
active
Recorded reasons:
| Category | agentic |
|---|---|
| Subcategory | independent re-score of Zapier AutomationBench (objectives minus guardrails) |
| Page status | active |
| Metric | share of objectives completed with no guardrail violation |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 657 |
| Publisher | Artificial Analysis |
AutomationBench-AA is Artificial Analysis's run of Zapier's AutomationBench on a private held-out split. The agent still has to complete SaaS workflows across simulated business apps by discovering REST APIs and writing correct state. AA does not use Zapier's strict all-assertions pass rate as the headline. It splits each assertion into an objective the agent must make true or a guardrail that already holds and must not be broken. The published score is the share of objectives completed, with the whole task zeroed if any guardrail fires. English text and structured tool calls; no screenshot browsing.
One run per task in Zapier's multi-turn AutomationBench environment, API toolset, 50-turn cap. Structured tool calls to search and execute REST endpoints. Programmatic end-state grading; no LLM judge.
Each row was checked against its source by a reviewer.
| Model | Score | Evidence date | Source kind | Link |
|---|---|---|---|---|
| GPT-6 Astra (max) | 68.5% | 2026-09-07 published | independent_evaluator | source |
| GLM-5.3 (max) | 62.2% | 2026-09-07 published | independent_evaluator | source |
No model card in ModelSpec reports this benchmark yet.