AutomationBench-AA

Artificial Analysis's independent run of Zapier's private AutomationBench split, scoring objective completion with zero credit after any guardrail break.

Also known as: AA AutomationBench, AutomationBench AA

active

This benchmark is in the default catalogue: its identity, protocol, current model coverage and dated results were verified by a reviewer who opened the sources.

Recorded reasons:

Categoryagentic
Subcategoryindependent re-score of Zapier AutomationBench (objectives minus guardrails)
Page statusactive
Metricshare of objectives completed with no guardrail violation
Directionhigher_is_better
Unit%
Dataset size657
PublisherArtificial Analysis

What it measures

AutomationBench-AA is Artificial Analysis's run of Zapier's AutomationBench on a private held-out split. The agent still has to complete SaaS workflows across simulated business apps by discovering REST APIs and writing correct state. AA does not use Zapier's strict all-assertions pass rate as the headline. It splits each assertion into an objective the agent must make true or a guardrail that already holds and must not be broken. The published score is the share of objectives completed, with the whole task zeroed if any guardrail fires. English text and structured tool calls; no screenshot browsing.

Task format

One run per task in Zapier's multi-turn AutomationBench environment, API toolset, 50-turn cap. Structured tool calls to search and execute REST endpoints. Programmatic end-state grading; no LLM judge.

Verified results

Each row was checked against its source by a reviewer.

ModelScoreEvidence dateSource kindLink
GPT-6 Astra (max)68.5%2026-09-07 publishedindependent_evaluatorsource
GLM-5.3 (max)62.2%2026-09-07 publishedindependent_evaluatorsource

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub