Signal Brief

Claude Opus 5 在 AA-Briefcase 基准上领先 Fable 5,成本更低

Artificial Analysis 发布评测,Claude Opus 5 在 AA-Briefcase 基准上以 1720 Elo 领先,比 Claude Fable 5 高 146 点,成本降低 20% 至 $17.79 每任务,但推理时间更长,max 设置平均 36.2 分钟。

twitter关注列表 Artificial Analysis (@ArtificialAnlys) 发布 2026-07-24 收录 2026-07-24 观察

一句话判断

新增数据点:Claude Opus 5 在 agentic 知识工作基准上超越 Fable 5 且成本更低,但推理时间显著增加,值得关注其实际应用中的效率权衡。

核心信息

Artificial Analysis 发布评测,Claude Opus 5 在 AA-Briefcase 基准上以 1720 Elo 领先,比 Claude Fable 5 高 146 点,成本降低 20% 至 $17.79 每任务,但推理时间更长,max 设置平均 36.2 分钟。

原始内容

Claude Opus 5 is the new leader on our agentic knowledge work benchmark, AA-Briefcase, outperforming Claude Fable 5 by nearly 150 Elo while reducing Cost per Task by 20% @AnthropicAI has released Claude Opus 5, the new leader on the Artificial Analysis Intelligence Index, and also the #1 model on our proprietary agentic knowledge work benchmark, AA-Briefcase. At max effort, Claude Opus 5 scores an AA-Briefcase Elo of 1720, 146 points ahead of Claude Fable 5 (1574). Its xhigh (1693) and high (1606) variants also outperform Fable 5 while being more cost efficient We benchmarked all five effort settings ahead of release. Claude Opus 5 (max) costs $17.79 per task, a 20% reduction from Claude Fable 5 ($22.30). The xhigh and high variants offer stronger cost-efficiency tradeoffs, both scoring above Fable 5 while costing 64% ($14.26) and 47% ($10.41) as much respectively AA-Briefcase is our proprietary benchmark for agentic knowledge work, evaluating models on realistic, private tasks spanning thousands of input files and requiring deliverables such as research reports, presentations, and spreadsheets. Performance is measured across correctness, analytical quality, and presentation quality, and combined into a single metric, the AA-Briefcase Elo Key AA-Briefcase results across all 5 Claude Opus 5 effort settings: ➤ The clear leader in agentic knowledge work: Claude Opus 5’s top three effort settings (max, xhigh, and high) take the top three positions on AA-Briefcase. Combined with Claude Fable 5, Sonnet 5 (max), and Opus 4.8 (max), Anthropic now holds a large majority of the top ten spots on the AA-Briefcase leaderboard. Claude Opus 5 scores 1470 Elo at medium effort, just behind GPT-5.6 Sol (max, 1505), and 1223 at low effort, just below GLM-5.2 (max, 1254) ➤ Leads in objective criteria and analytical quality, but not presentation: AA-Briefcase measures performance across objective rubric criteria, analytical quality, and presentation quality. Claude Opus 5’s gains are driven primarily by rubric pass rate and analytical quality. At max effort it achieves an Analytical Quality Elo of 2016, nearly 300 Elo ahead of Claude Fable 5. Compared to Opus 4.8, presentation quality sees a smaller improvement, with a Presentation Elo of 1628, still ~40 Elo behind GPT-5.6 Sol (max, 1666) ➤ Half the cost for Fable-level knowledge work capabilities: Claude Opus 5 offers a wide range of intelligence-cost tradeoffs across its five effort settings. Max effort costs $17.79 per task, around 20% less than Claude Fable 5 ($22.30) while achieving a higher AA-Briefcase Elo. Notably, Claude Opus 5 with high effort outperforms Claude Fable 5 by 32 Elo while costing $10.41 per task, less than half the price. Claude Opus 5 retains Opus 4.8’s pricing of $5/$25 per million input/output tokens, with a 25% premium for cache writes and a 90% discount for cache hits ➤ Over 25 minutes per task for frontier performance: Claude Opus 5’s top three effort settings all average more than 25 minutes per AA-Briefcase task (36.2, 34.3, and 25.7 minutes for max, xhigh, and high). At max effort this is roughly 50% longer than Opus 4.8 (24.1 minutes), primarily due to an increase in the number of turns. Claude Opus 5 averages 103, 91, and 76 turns per task across its top three effort levels, compared to 55 for Opus 4.8 (max) ![photo](https://pbs.twimg.com/media/HOBjK6cbIAA2Yph.jpg) Artificial Analysis (@ArtificialAnlys): For full AA-Briefcase results, see https://t.co/RgkI2BmI6R Artificial Analysis (@ArtificialAnlys): Claude Opus 5’s top three effort settings all average more than 25 minutes per AA-Briefcase task (36.2, 34.3, and 25.7 minutes for max, xhigh, and high). At max effort this is roughly 50% longer than Opus 4.8 (24.1 minutes), primarily due to an increase in the number of turns. Claude Opus 5 averages 103, 91, and 76 turns per task across its top three effort levels, compared to 55 for Opus 4.8 (max)

相关动态

01

Atomic Agent GAIA基准胜Hermes

Atomic Agent在GAIA Level 1基准测试中以69.8%正确率超越Hermes的58.5%,速度快1.6倍(3h12m vs 5h10m),使用相同4-bit Qwen-3.6-35B模型和Apple M4 Max硬件,开源MIT许可。

twitter关注列表2026-07-24#开源#评测#技术突破
观察
03

Grok 集成 Google Workspace

Grok 已集成到 Google Workspace,通过侧边栏插件在 Sheets、Slides、Docs 中提供数据解释、公式书写、演示生成、文档起草等功能,无需频繁切换标签页,组织可批量部署。

twitter关注列表2026-07-24#产品发布#大模型#AI
观察