Signal Brief

行业 AI 模型能力指数发布

Artificial Analysis 发布六个新的行业 AI 模型能力指数,覆盖金融、法律、医疗、战略运营、工程和经济领域,基于 O*NET 任务分类并独立运行基准评估。结果显示 Claude Fable 5 在所有指数中领先,GLM-5.2 在开放权重模型中领先五个行业指数,DeepSeek V...

twitter关注列表 Artificial Analysis (@ArtificialAnlys) 发布 2026-07-07 收录 2026-07-07 观察

一句话判断

新增跨领域模型能力量化指标,独立评测方法论公开,适合用于模型选型参考。

核心信息

Artificial Analysis 发布六个新的行业 AI 模型能力指数,覆盖金融、法律、医疗、战略运营、工程和经济领域,基于 O*NET 任务分类并独立运行基准评估。结果显示 Claude Fable 5 在所有指数中领先,GLM-5.2 在开放权重模型中领先五个行业指数,DeepSeek V4 Flash 以极低成本完成所有行业任务。

原始内容

Introducing six new Artificial Analysis Capability Indices for comparing model capabilities across key industry domains The new industry indices cover Finance & Accounting, Legal, Healthcare & Medical, Strategy & Ops, Engineering, and Economics. We aim to capture the common capabilities required across knowledge work domains and evaluate how well current models meet those needs. Each index is grounded in common tasks from O*NET occupational classifications. Tasks range from financial modeling, to legal research and contract review, to clinical decision support and patient documentation. We derive capabilities from each task, select the benchmarks that best represent the work, and weight by how often each capability appears across the domain. This means rethinking the Artificial Analysis benchmark suite for each domain and slicing evaluations to relevant domain tasks. Every component benchmark is run independently by Artificial Analysis. The industry indices join the existing skill-based Agentic and Coding indices, which measure capabilities that cut across every domain. Key Results ➤ Leading models: Claude Fable 5 (with Opus 4.8 fallback) leads all eight indices, with Claude Opus 4.8 (max) in second on six of eight Capability Indices and GPT-5.5 (xhigh) on two. Below the top two, rankings reshuffle substantially by domain between Gemini 3.5 Flash, Gemini 3.1 Pro Preview, GPT-5.5 (xhigh), Claude Sonnet 5 (max), and GLM-5.2 (max). ➤ Open weights leading models: Among open weights models, GLM-5.2 (max) leads on five of the six industry indices, ranking as high as fifth overall on the Artificial Analysis Engineering Index (53), within 2 points of Claude Sonnet 5 (max, 55) and GPT-5.5 (xhigh, 55). DeepSeek V4 Pro (max, 38) takes the open weights lead on Artificial Analysis Strategy & Ops Index. ➤ Cost efficiency: DeepSeek V4 Flash (max) completes tasks for <$0.04 across all six indices while scoring mid-pack, and GLM-5.2 (max) leads open weights score with a Cost per Task of $0.26 to $0.58. Frontier capability comes at a steep premium: on the Artificial Analysis Strategy & Ops Index, Claude Fable 5 (with Opus 4.8 fallback, $3.48) scores 12 points above DeepSeek V4 Pro (max, $0.03) at over 100x the Cost per Task. ➤ Time per Task: Time per Task spreads roughly 15x within each index, from 1.1 minutes for Nova 2.0 Pro Preview (medium) to 16.7 minutes for Claude Sonnet 5 (max). Speed shows a similar frontier to cost: on the Artificial Analysis Legal Index, Gemini 3.1 Pro Preview (0.8 minutes) completes tasks ~7x faster than Claude Fable 5 (with Opus 4.8 fallback, 5.4 minutes), while scoring within 11 points. ![photo](https://pbs.twimg.com/media/HMlTMY7aQAA2wWR.jpg) Artificial Analysis (@ArtificialAnlys): Explore all Capability Indices, with full weightings and component benchmarks, on the Artificial Analysis website: https://t.co/0jSq4HtZAr Finance & Accounting: https://t.co/cu82nCqSOe Strategy & Ops: https://t.co/Z8Y8NBIYvx Legal: https://t.co/qbjxPHcaLs Healthcare & Medical: https://t.co/mCWo4lRzlF Engineering: https://t.co/wzjAfywmxG Economics: https://t.co/qjCNODzuig As always, our full methodology is public. See how evaluations are run, scored, and weighted on each index: https://t.co/edIbUixO8U

相关动态

01

Kimi K3 在 SpreadsheetBench 2 上排名第一

Kimi K3 在 SpreadsheetBench 2 基准测试中排名第一,超越 Claude Fable 5,完成 34.8% 的 workflow 任务。该基准测试涉及 321 个专家策划任务,平均每个任务包含 11.8 个工作表和 593.5 个单元格更改,聚焦于完整的电子表格工作簿执行。

twitter关注列表2026-07-18#模型#技术突破#评测
观察
02

AI 代码生成大势所趋

Greg Isenberg 发文指出,相比一年前的手写代码实践,如今大多数工程代码已由 AI 生成,标志着编程范式的根本转变。他援引了 Google 75% 新代码由 AI 生成、Anthropic 90%+ 代码由 Claude 编写、GitClear 代码重复率上升 81% 复用率下降 70% 等具体数据,并引用 Dario Amod...

twitter关注列表2026-07-18#技术突破#行业动态#分析
观察
03

Kimi K3 编码代理评测

Artificial Analysis发布Kimi K3在编码代理指数上的评测结果:得分57,排名第5,性能与GPT-5.6 Terra和GPT-5.5持平(57),超过Opus 4.8(55),略低于Grok 4.5(58)和Fable 5(59);成本平均3.18美元/任务,比GPT-5.6 Sol便宜55%。

twitter关注列表2026-07-17#AI模型#评测
观察
04

Runway Agent 在第三方评测中全面领先

Physion Labs 发布 Physion-Arc 1.0 基准测试,对 Runway、Luma、MiniMax、Kling、Utopia 和 TapNow 六款 AI 视频代理进行独立人类评估,使用 30 个电影提示和 16 个指标。Runway Agent 2.0 在叙事连贯性、电影语言和制作质量三项核心维度全部排名第一。

twitter关注列表2026-07-17#评测#多模态#模型
观察
05

Grok 4.5成本效率对比评测

Artificial Analysis数据显示,Grok 4.5每次任务成本仅0.31美元,而Claude Fable 5为2.75美元、Claude Opus 4.8为1.80美元、GPT-5.6 Sol为1.04美元、Kimi K3为0.95美元,Grok 4.5成本分别低约9倍、6倍、3倍和3倍;在FrontierSWE上,Grok...

twitter关注列表2026-07-17#AI#模型#评测
观察