Signal Brief

AutomationBench-AA评测

Artificial Analysis与Zapier合作发布AutomationBench-AA基准测试,评估AI agent在真实SaaS工作流中的表现,涵盖Finance、HR等6个领域的657个任务,跨40个模拟应用。Claude Fable 5以48.6%得分领先,Opus 4.8为48.5...

twitter关注列表 Artificial Analysis (@ArtificialAnlys) 发布 2026-07-06 收录 2026-07-06 观察

一句话判断

该基准提供了多模型在复杂工作流中的细粒度表现数据,包括防护栏违规和成本效率,值得点开原文查看完整榜单和方法论。

核心信息

Artificial Analysis与Zapier合作发布AutomationBench-AA基准测试,评估AI agent在真实SaaS工作流中的表现,涵盖Finance、HR等6个领域的657个任务,跨40个模拟应用。Claude Fable 5以48.6%得分领先,Opus 4.8为48.5%,Gemini 3.5 Flash为42.6%,GPT-5.5为42.1%。Gemini 3.5 Flash每任务成本0.49美元,GPT-5.5为1.32美元。所有模型均违反商业规则,Gemini 3.5 Flash每任务防护栏违规0.46次,比例最优。

原始内容

Announcing AutomationBench-AA, our independent leaderboard for Zapier’s AutomationBench, testing whether AI agents can automate real SaaS workflows while adhering to business rules We partnered with @zapier to run AutomationBench-AA on their private benchmark subset. This benchmark is a complex agentic workflow automation test across simulated SaaS applications. Models must complete 657 tasks spanning Finance, HR, Marketing, Operations, Sales, and Support, working across 40 simulated app environments including Gmail, Google Sheets, Slack, Salesforce, Zendesk, Jira, and HubSpot. Unlike Zapier’s hosted leaderboard, the headline score for AutomationBench-AA shows the share of objectives a model completes without violating any guardrails. Claude Fable 5 and Opus 4.8 from @AnthropicAI lead with scores of 48.6% and 48.5%, followed by @GoogleDeepMind's Gemini 3.5 Flash at 42.6% and @OpenAI's GPT-5.5 (xhigh) at 42.1%. With Anthropic’s new classifier, Fable 5 fell back to Opus on ~18% of tasks. Key elements of AutomationBench: ➤ Real workflow patterns, simulated environments: Tasks are drawn from real workflow patterns on Zapier and run in simulated SaaS environments, where a single task may span a range of applications like CRM, email, calendar, and messaging platforms. ➤ Autonomous API discovery: Models interact with each app through REST APIs, discovering the endpoints they need through structured tool calls and navigating environments with irrelevant and sometimes misleading records. ➤ Objectives and guardrails: Models are scored against nearly 12,000 assertions Zapier built to test that the model completed the task correctly in full. Each assertion is classified as either an objective the agent must achieve, or a guardrail that already passes initially and must not be broken. ➤ Programmatic environment grading: Tasks are graded solely on whether the correct data ended up in the right systems, with deterministic checks against the environment. Each task runs once with a 50-turn cap. Key results for AutomationBench-AA: ➤ Claude Fable 5 (max) leads at 48.6% but falls back to Opus 4.8 in ~18% of tasks. It completes 73% of task objectives, with the fallback behavior likely explaining the limited uplift compared to Opus. ➤ Every model breaks business rules: Guardrail violations range from 0.46 per task (Gemini 3.5 Flash) to 1.26 (Qwen3.7 Plus). Gemini 3.5 Flash completes 15.0 objectives per guardrail violation, the best ratio of any model, ahead of Claude Opus 4.8 (max, 13.5). ➤ Gemini 3.5 Flash performs well for its price: At 42.6% and $0.49 per task, it effectively matches GPT-5.5 (xhigh, 42.1%, $1.32 per task) at ~37% of the cost. ➤ GLM-5.2 (max) from @Zai_org is the leading open weights model at 27.8%. This places the open weights frontier ~10 points behind Gemini 3.1 Pro Preview, and with substantially higher guardrail violations per task. ➤ Finance workflow tasks are the most difficult to automate today: across the models we evaluated at launch, agents complete roughly half the proportion of objectives on Finance tasks, compared to Support and Operations tasks. We would like to thank Zapier and the benchmark authors for their great work developing this evaluation for important SaaS workflows, and appreciate their collaboration in launching AutomationBench on Artificial Analysis! ![photo](https://pbs.twimg.com/media/HMkBhrbaQAAUuVd.jpg) Artificial Analysis (@ArtificialAnlys): Working styles differ sharply on AutomationBench. GPT-5.5 (xhigh) takes a more action-intensive approach, averaging 49 tool calls across 25 turns per task, while Claude Opus 4.8 (max) is more deliberate: 35 tool calls packed into just 14 turns, with fewer guardrail violations (0.55 vs 0.66 per task). Grok 4.3 (high) takes the fewest turns (13), but does not perform as strongly as models which persist longer to complete tasks, consistent with declaring tasks complete prematurely rather than finishing them efficiently. Artificial Analysis (@ArtificialAnlys): For more information see full results of AutomationBench-AA on our website: https://t.co/cm8vKYxq0N Zapier leaderboard: https://t.co/gm3ML8O0kW ArXiv paper: https://t.co/EyvuwlaiT6 AutomationBench on GitHub: https://t.co/UHmewKhJ2K

相关动态

01

Kimi K3 在 SpreadsheetBench 2 上排名第一

Kimi K3 在 SpreadsheetBench 2 基准测试中排名第一,超越 Claude Fable 5,完成 34.8% 的 workflow 任务。该基准测试涉及 321 个专家策划任务,平均每个任务包含 11.8 个工作表和 593.5 个单元格更改,聚焦于完整的电子表格工作簿执行。

twitter关注列表2026-07-18#模型#技术突破#评测
观察
02

Mind Lab 开源长文本强化学习项目

Mind Lab 开源了一个长文本强化学习(RL)项目,支持 2M tokens 的长上下文。相比以往需要数千个 GPU 的 1M 上下文研究,该项目仅需 8 个 GPU 即可完成,大幅降低了超长上下文研究的算力门槛。

twitter关注列表2026-07-17#开源#技术突破#模型
观察
03

Kimi K3非对称竞争分析

Kimi K3在Arena前端/WebDev人类盲测中以1679 Elo排名第一,领先Fable 5(1631)和GPT-5.6 Sol(1618),但在综合智能指数中仅排第四(57.1分),落后Fable 5(59.9)和GPT-5.6 Sol(58.9)。K3定价15美元/百万输出tokens,远低于Fable 5的50美元,并计划7...

twitter关注列表2026-07-18#AI#模型#技术突破
观察
04

AI 代码生成大势所趋

Greg Isenberg 发文指出,相比一年前的手写代码实践,如今大多数工程代码已由 AI 生成,标志着编程范式的根本转变。他援引了 Google 75% 新代码由 AI 生成、Anthropic 90%+ 代码由 Claude 编写、GitClear 代码重复率上升 81% 复用率下降 70% 等具体数据,并引用 Dario Amod...

twitter关注列表2026-07-18#技术突破#行业动态#分析
观察
05

开源模型长期网络能力差距缩小至4-7个月

长期网络能力方面,领先开源模型落后闭源前沿从2025年大部分时间的6-10个月缩短至4-7个月。GLM-5.2匹配了约4个月前发布的闭源模型。AI安全研究所的32步“The Last Ones”任务中,GPT-5.6 Sol在10次尝试中完成7次,每次预算1亿token,且性能随推理token增加而提升。

twitter关注列表2026-07-17#模型#技术突破#AI安全
观察