Signal Brief

六项行业能力基准发布

Artificial Analysis 发布六项新的行业能力基准,覆盖金融、法律、医疗、策略与运营、工程和经济六个领域。基准结果显示,Claude Fable 5(搭配 Opus 4.8 回落)在全部八项指数中排名第一,Claude Opus 4.8(最大版)在六项指数中位列第二,GPT-5.5(x...

twitter关注列表 Chubby♨️ (@kimmonismus) 发布 2026-07-07 收录 2026-07-07 观察

一句话判断

该基准集以 O*NET 职业分类为依据,按任务频率加权,提供跨域能力的可比较度量。

核心信息

Artificial Analysis 发布六项新的行业能力基准,覆盖金融、法律、医疗、策略与运营、工程和经济六个领域。基准结果显示,Claude Fable 5(搭配 Opus 4.8 回落)在全部八项指数中排名第一,Claude Opus 4.8(最大版)在六项指数中位列第二,GPT-5.5(xhigh)在两项指数中位列第二。在开源模型中,GLM-5.2(最大版)在五项行业基准中领先,工程指数排名第五(得分 53),仅低于 Claude Sonnet 5(最大版,55)和 GPT-5.5(xhigh,55)两分;成本方面,DeepSeek V4 Flash 单任务费用低于 0.04 美元,而 Claude Fable 5 单任务费用达 3.48 美元,速度上 Nova 2.0 Pro Preview 中等版每任务 1.1 分钟,Claude Sonnet 5 最大版达 16.7 分钟。

原始内容

tl;dr: Fable 5 basically mocks every other model and tops every benchmark. Curious to see whether the six new benchmarks will crown a new leader after GPT-5.6. https://t.co/fsMk3NKif8 ![photo](https://pbs.twimg.com/media/HMnq3ydWIAAED7y.png) > **引用原帖 Artificial Analysis (@ArtificialAnlys):** > Introducing six new Artificial Analysis Capability Indices for comparing model capabilities across key industry domains > The new industry indices cover Finance & Accounting, Legal, Healthcare & Medical, Strategy & Ops, Engineering, and Economics. We aim to capture the common capabilities required across knowledge work domains and evaluate how well current models meet those needs. > Each index is grounded in common tasks from O*NET occupational classifications. Tasks range from financial modeling, to legal research and contract review, to clinical decision support and patient documentation. We derive capabilities from each task, select the benchmarks that best represent the work, and weight by how often each capability appears across the domain. This means rethinking the Artificial Analysis benchmark suite for each domain and slicing evaluations to relevant domain tasks. Every component benchmark is run independently by Artificial Analysis. > The industry indices join the existing skill-based Agentic and Coding indices, which measure capabilities that cut across every domain. > Key Results > ➤ Leading models: Claude Fable 5 (with Opus 4.8 fallback) leads all eight indices, with Claude Opus 4.8 (max) in second on six of eight Capability Indices and GPT-5.5 (xhigh) on two. Below the top two, rankings reshuffle substantially by domain between Gemini 3.5 Flash, Gemini 3.1 Pro Preview, GPT-5.5 (xhigh), Claude Sonnet 5 (max), and GLM-5.2 (max). > ➤ Open weights leading models: Among open weights models, GLM-5.2 (max) leads on five of the six industry indices, ranking as high as fifth overall on the Artificial Analysis Engineering Index (53), within 2 points of Claude Sonnet 5 (max, 55) and GPT-5.5 (xhigh, 55). DeepSeek V4 Pro (max, 38) takes the open weights lead on Artificial Analysis Strategy & Ops Index. > ➤ Cost efficiency: DeepSeek V4 Flash (max) completes tasks for <$0.04 across all six indices while scoring mid-pack, and GLM-5.2 (max) leads open weights score with a Cost per Task of $0.26 to $0.58. Frontier capability comes at a steep premium: on the Artificial Analysis Strategy & Ops Index, Claude Fable 5 (with Opus 4.8 fallback, $3.48) scores 12 points above DeepSeek V4 Pro (max, $0.03) at over 100x the Cost per Task. > ➤ Time per Task: Time per Task spreads roughly 15x within each index, from 1.1 minutes for Nova 2.0 Pro Preview (medium) to 16.7 minutes for Claude Sonnet 5 (max). Speed shows a similar frontier to cost: on the Artificial Analysis Legal Index, Gemini 3.1 Pro Preview (0.8 minutes) completes tasks ~7x faster than Claude Fable 5 (with Opus 4.8 fallback, 5.4 minutes), while scoring within 11 points. > https://x.com/ArtificialAnlys/status/2074299714699469221

相关动态

02

AI 代码生成大势所趋

Greg Isenberg 发文指出,相比一年前的手写代码实践,如今大多数工程代码已由 AI 生成,标志着编程范式的根本转变。他援引了 Google 75% 新代码由 AI 生成、Anthropic 90%+ 代码由 Claude 编写、GitClear 代码重复率上升 81% 复用率下降 70% 等具体数据,并引用 Dario Amod...

twitter关注列表2026-07-18#技术突破#行业动态#分析
观察
03

AI可能替代低质量人工代码工作

作者和Tobi Lutke分析AI代码替代趋势,指出大量低效人力资源可能被AI替代,核心阻力是系统集成问题。认为替代阻力不是技术瓶颈,而是与现有人工工作质量和流程的适配问题。

twitter关注列表2026-07-17#AI#技术更新#分析
观察
04

习近平在2026世界人工智能大会反对美限制 推开源治理承诺5000名额

中国国家主席习近平在上海2026世界人工智能大会致辞,反对美国主导的AI技术限制,倡导开源AI与全球共治,警告以国家安全为由泛化限制将让强国垄断技术获取权。中国承诺未来5年为发展中国家伙伴提供5000个AI研究、培训与合作名额,将开源AI定位为产业政策与外交工具双重角色。

twitter关注列表2026-07-17#政策#行业动态#开源
值得跟进
05

习近平在世界AI大会阐述AI全球治理愿景

习近平在首次出席世界人工智能大会时提出中国对全球AI秩序的愿景,倡导开源AI以促进开放共赢,反对美国以国家安全为名限制技术共享,并宣布未来五年为发展中国家提供5000个AI培训机会,与东盟、阿盟、非盟等建立合作中心。

twitter关注列表2026-07-17#AI#政策#行业动态
观察