Signal Brief

Artificial Analysis发布Harvey LAB-AA法律基准测试

Artificial Analysis重新实现Harvey的法律代理基准Harvey LAB-AA,使用由Harvey团队构建的120个.scenario测试了24个法律领域的模型。Claude Fable 5(AnthropicAI)以14.2%的全通过率遥遥领先,其次是Claude Opus 4...

twitter关注列表 Artificial Analysis (@ArtificialAnlys) 发布 2026-07-07 收录 2026-07-07 观察

一句话判断

该基准首次量化法律领域AI的完整任务完成率,并揭示成本与表现的巨大差距,值得深入评估。

核心信息

Artificial Analysis重新实现Harvey的法律代理基准Harvey LAB-AA,使用由Harvey团队构建的120个.scenario测试了24个法律领域的模型。Claude Fable 5(AnthropicAI)以14.2%的全通过率遥遥领先,其次是Claude Opus 4.8和GLM-5.2(Zai_org)各7.5%。测试包括每个任务的二进制评判标准,最高模型在单个评判标准上超过90%通过率,但大多数模型在任何任务上无法完成全部要求。成本差距显著,Claude Fable 5约19美元/任务,Gemini 3.1 Flash‑Lite约0.02美元/任务。

原始内容

After our announcement last month, Artificial Analysis is now launching Harvey LAB-AA (Legal Agent Benchmark), our implementation of Harvey's new agentic legal benchmark that evaluates language models on real-world legal work across 24 practice areas Harvey LAB-AA tests models on a private set of 120 legal tasks built by the team at @harvey. The tasks span 24 practice areas from corporate M&A and capital markets to tax, litigation, and bankruptcy. Models work to create the legal outputs specified in the tasks, and each task is graded against a rubric of binary criteria. The primary metric we present is the all-pass rate: the share of tasks where all criteria in the rubric are satisfied, reflecting the high standard of real-world professional legal deliverables. Claude Fable 5 (max, with fallback) from @AnthropicAI leads Harvey LAB-AA with a 14.2% all-pass rate, after falling back to Opus 4.8 in only 1 task. This is almost double the scores of the next best models Claude Opus 4.8 (max) and GLM-5.2 (max) from @Zai_org, which tie at 7.5%. Key takeaways from Harvey LAB-AA: ➤ Frontier legal work is far from solved: At launch, most models pass a majority of individual criteria but very few fully satisfy the requirements of any given task. The best model, Claude Fable 5, fully satisfies rubrics on just 14.2% of tasks, leaving ~86% of professional legal deliverables incomplete. Claude Opus 4.8 (max) and GLM-5.2 (max) follow at 7.5%, MiniMax-M3 at 6.7%, and Claude Sonnet 5 at 5.0%, ahead of GPT-5.5 (xhigh) from @OpenAI and Claude Sonnet 4.6 (max), which both score 4.2%. ➤ Models can pass many requirements of legal tasks, but rarely all of them: the leading models pass >90% of individual rubric criteria, but 13 of the 28 evaluated models fully pass 0 tasks. ➤ The top open weights model scores just over half the frontier leader: GLM-5.2 (max) ties Claude Opus 4.8 for second with a 7.5% all-pass rate and criteria pass of 91.0% vs. 91.1% respectively, both now behind Claude Fable 5 (14.2%). GLM-5.2 reaches that at ~6% of Fable 5's cost per task (~$1 vs. ~$19). ➤ Cost per task spans ~950x: the most expensive model, Claude Fable 5, costs ~$19 per task, while Gemini 3.1 Flash-Lite passes 31.1% of criteria for ~$0.02 per task. ![photo](https://pbs.twimg.com/media/HMo9646bQAAJNQC.jpg) Artificial Analysis (@ArtificialAnlys): To see Harvey’s original LAB announcement, please visit https://t.co/LiSeheHGSY Explore the full Harvey LAB-AA results at https://t.co/vSNENODbYR Artificial Analysis (@ArtificialAnlys): Harvey LAB-AA is our independent reimplementation of Harvey’s evaluation, and there are several key differences to the original version: ➤ Models are run on our Stirrup agent harness, enabling features such as context compaction rather than failure when reaching context limits, with simplified Artificial Analysis-authored agent and judge prompts ➤ We do not include Harvey's custom tools and document-generation skill scripts (e.g., pptx, docx), instead leveraging a simple code execution tool to reflect raw model ability ➤ Deliverables must match the exact filename specified, rather than fuzzy and LLM-based matching when models produce incorrect filenames ➤ Grading uses a single Gemini 3.1 Pro judge, tested to be well-calibrated against a frontier panel For more information on the Harvey LAB-AA methodology, please visit https://t.co/bSGfKE8tHX

相关动态

02

Kimi K3 在 SpreadsheetBench 2 上排名第一

Kimi K3 在 SpreadsheetBench 2 基准测试中排名第一,超越 Claude Fable 5,完成 34.8% 的 workflow 任务。该基准测试涉及 321 个专家策划任务,平均每个任务包含 11.8 个工作表和 593.5 个单元格更改,聚焦于完整的电子表格工作簿执行。

twitter关注列表2026-07-18#模型#技术突破#评测
观察
03

AI 代码生成大势所趋

Greg Isenberg 发文指出,相比一年前的手写代码实践,如今大多数工程代码已由 AI 生成,标志着编程范式的根本转变。他援引了 Google 75% 新代码由 AI 生成、Anthropic 90%+ 代码由 Claude 编写、GitClear 代码重复率上升 81% 复用率下降 70% 等具体数据,并引用 Dario Amod...

twitter关注列表2026-07-18#技术突破#行业动态#分析
观察
04

BestBlogs 早报 · 07-18

月之暗面发布Kimi K3,2.8万亿参数,896选16的Stable LatentMoE,上下文100万token,接近Fable-5但仍落后最强闭源模型,完整权重7月27日前开源;VentureBeat调查显示54%企业已发生AI代理安全事件;xAI开源Grok Build(84万行Rust代码)并残留上传用户代码痕迹;Cursor评...

twitter关注列表2026-07-17#AI#模型发布#开源
观察
05

Kimi K3 编码代理评测

Artificial Analysis发布Kimi K3在编码代理指数上的评测结果:得分57,排名第5,性能与GPT-5.6 Terra和GPT-5.5持平(57),超过Opus 4.8(55),略低于Grok 4.5(58)和Fable 5(59);成本平均3.18美元/任务,比GPT-5.6 Sol便宜55%。

twitter关注列表2026-07-17#AI模型#评测
观察