Benchmarks · Agents, tools & computer use
AutomationBench
Can an AI agent carry out a real business workflow across multiple SaaS applications — finding the right API endpoints itself, following a company's own layered business rules, and getting the right data to the right system — rather than completing a single, well-specified task?
ZapierReleased 21 April 2026Live
Business automation is usually pitched as connecting one app to another. AutomationBench, built by Zapier researchers Daniel Shepard and Robin Salimans, tests something closer to what that actually requires in practice: an agent dropped into a simulated company with dozens of SaaS tools, asked to complete a workflow — routing a lead, processing an expense, closing a support ticket — without being told which tool exposes which API. The agent has to find the right endpoints itself, follow the business’s own layered approval rules, and avoid being thrown off by the irrelevant or misleading records that fill a real system.
Drawing on patterns from Zapier’s own automation platform, the benchmark spans six domains — Sales, Marketing, Operations, Support, Finance and HR — with around 1,200 tasks split between a public set and a larger, harder private set reserved for the official leaderboard. Scoring is entirely programmatic: a task counts as solved only if the correct data ends up in the correct place, with no partial credit for a plausible-looking but incorrect action.
The results at release were a sharp contrast to the near-saturated benchmarks common elsewhere in agentic evaluation: every frontier model scored under 10%, with Claude Opus 4.7 leading at 9.9%, Gemini 3.1 Pro close behind at 9.6%, and GPT-5.4 at 7.6%. Scores rose over the following months — Claude Fable 5 reached 17.4% in June and Kimi K3 reported 30.8% on the public subset in July — but even the leader still failed roughly two-thirds of tasks. That gap between models’ fluency on short, well-specified coding and reasoning tasks and their fragility on multi-step, real-world business workflows is precisely what AutomationBench was built to expose.
The set
Around 1,200 tasks (600+ public, 600+ held-back private) spanning six business domains — Sales, Marketing, Operations, Support and HR — built from real workflow patterns observed on Zapier's automation platform across 47 simulated tools. Agents must discover relevant API endpoints themselves, respect business-policy constraints, and work in environments seeded with irrelevant or misleading records; scoring is fully programmatic, checking only whether the correct end state was reached.
Example
There's a scheduling conflict on February 20, 2026 at 2:00 PM — a Zoom meeting and a Google Calendar event overlap. Check the meeting priority policy in the spreadsheet to determine which one wins, then reschedule the loser by prepending [RESCHEDULED] to its topic/title. Post a summary to #ops-updates on Slack noting which meeting won and which was rescheduled, including both the Zoom meeting ID and Calendar event ID.arxiv.org
Where it stands
At the April 2026 release even frontier models scored under 10% (Claude Opus 4.7 led at 9.9%). Scores have since climbed but stayed low: Kimi K3 reported 30.8% on the public subset in July 2026 and GPT-6 Astra reported 41.4% in September 2026, still leaving well over half of business workflows unsolved — one of the least-saturated agentic benchmarks of its period.
How the top score changed hands
Best reported score at each point the lead changed hands. Tap a marker for the entry.
- April 2026Claude Opus 4.79.9%Top score at release, ahead of Gemini 3.1 Pro (9.6%) and GPT-5.4 high (7.6%) — every frontier model under 10%.
- June 2026Claude Fable 517.4%Company-reported; the frontier had roughly doubled in two months.
- July 2026Kimi K330.8%On the 600-task public subset; the mid-2026 high.
- September 2026GPT-6 Astra41.4%GPT-6 Astra launch table; the highest figure recorded here, still short of solving the benchmark.
Current best: GPT-6 Astra — 41.4% AutomationBench (business workflows); GPT-6 Astra launch table (Sol 18.1%, Fable 5.1 31.4%, Opus 5 26.9%). Company-reported; still fails nearly three-fifths of tasks.
In the timeline · 3 entries
Anthropic releases Claude Opus 5.5
Anthropic said Opus 5.5 matches Fable 5.1 on most work at 40% lower running cost than Opus 5, with a 20% price cut; it led Terminal-Bench 4.0 and Artificial Analysis's GDPval-AA board.
Models & capabilities · Benchmarks & progress
OpenAI releases GPT-6 Astra
OpenAI's flagship is its first model rated 'Critical' for cyber capability, and its launch is shadowed by disclosures that a 'recurrent depth' technique makes Astra's reasoning harder to monitor.
Models & capabilities · Safety & alignment · Security & misuse
Anthropic releases Claude Fable 5.1 and Claude Mythos 5.1
Anthropic's launch table put Fable 5.1 ahead of its own Fable 5 and Opus 5 and OpenAI's GPT-5.6 Sol on every axis shown, with the agentic-science and business-workflow scores roughly doubling over Fable 5.
Models & capabilities · Benchmarks & progress