Signal Brief

8天内四款前沿模型发布评测

Artificial Analysis 发布评测,8天内 SpaceXAI 的 Grok 4.5(54分)、OpenAI 的 GPT-5.6 Sol/Terra/Luna(最高59/55/51)、Meta 的 Muse Spark 1.1(51分)和 Moonshot AI 的 Kimi K3(57...

twitter关注列表 Artificial Analysis (@ArtificialAnlys) 发布 2026-07-17 收录 2026-07-17 值得跟进

一句话判断

提供了各模型具体得分、价格和性能对比,以及前沿格局从2家扩展到6家的关键变化,值得阅读原文获取详细数据。

核心信息

Artificial Analysis 发布评测,8天内 SpaceXAI 的 Grok 4.5(54分)、OpenAI 的 GPT-5.6 Sol/Terra/Luna(最高59/55/51)、Meta 的 Muse Spark 1.1(51分)和 Moonshot AI 的 Kimi K3(57分)相继发布,前沿模型实验室从6月初的2家增至6家。Kimi K3 以57分位列第三,在知识工作基准上仅次于 Claude Fable 5,且成本仅为后者一半。前沿智能价格在8天内下降2-3倍。

原始内容

The frontier has opened up: four frontier launches in eight days - Grok 4.5, GPT-5.6, Muse Spark 1.1, and yesterday Kimi K3; six labs now have a model scoring >50 on the Artificial Analysis Intelligence Index, up from two in early June @SpaceXAI's Grok 4.5 (high, Artificial Analysis Intelligence Index Score: 54) landed July 8, @OpenAI's GPT-5.6 Sol, Terra, and Luna (max: 59, 55, 51) and @AIatMeta's Muse Spark 1.1 (xhigh, 51) followed a day later, and @Kimi_Moonshot's Kimi K3 launched yesterday at 57 - third overall, ahead of Claude Opus 4.8 (max, 56). The top three models on the Index now come from three different labs and span just three points. Four of the ten highest-scoring models launched since July 8, and six of ten since early June. The one thing that did not move is #1: Claude Fable 5 (max, 60) has held the top spot since June 9, but its lead has narrowed from four points to one, and the price of the intelligence beneath it collapsed. Congratulations to @elonmusk, @sama, @finkd, and the teams at all four labs on a remarkable eight days. Key Takeaways: ➤ The frontier went from two labs to six in six weeks. Until June, only Anthropic and OpenAI had fielded a model at 51 or above. GLM-5.2 (max) brought Z AI in mid-June; last week added SpaceXAI and Meta; today Kimi K3 makes Moonshot AI the sixth as they enter at 57 ➤ Kimi K3 debuts at #3 with agentic and knowledge work scores behind only the top two. K3 scores 1668 Elo on GDPval-AA v2, third behind Claude Fable 5 (max, 1760) and GPT-5.6 Sol (max, 1748). On AA-Briefcase, our benchmark of long-horizon knowledge work, it enters at #2 with 1547 Elo - behind only Claude Fable 5 (max, 1583) and ahead of GPT-5.6 Sol (max, 1495) - with an Analytical Quality Elo (1760) effectively tied with Fable 5 (1764). At $0.94 per Intelligence Index task on its $3/$15 pricing, it delivers comparable intelligence to Claude Opus 4.8 (max, $1.80) at roughly half the cost per task ➤ Near-frontier intelligence got 2-3x cheaper in eight days. GPT-5.6 Sol (max) delivers one point below Claude Fable 5 (max) at $1.04 per Intelligence Index task vs $2.75. Grok 4.5 (high) delivers 54 at $0.31, under a third of GPT-5.5 (xhigh, $0.99). At 51, GPT-5.6 Luna (max, $0.21) and Muse Spark 1.1 (xhigh, $0.26) undercut GLM-5.2 (max, $0.32), the cheapest at that level a week earlier ![photo](https://pbs.twimg.com/media/HNcf2upbMAALvLu.jpg) Artificial Analysis (@ArtificialAnlys): See more on Artificial Analysis: https://t.co/PQCRupCPta

相关动态

01

Kimi K3 在 SpreadsheetBench 2 上排名第一

Kimi K3 在 SpreadsheetBench 2 基准测试中排名第一,超越 Claude Fable 5,完成 34.8% 的 workflow 任务。该基准测试涉及 321 个专家策划任务,平均每个任务包含 11.8 个工作表和 593.5 个单元格更改,聚焦于完整的电子表格工作簿执行。

twitter关注列表2026-07-18#模型#技术突破#评测
观察
02

Mind Lab 开源长文本强化学习项目

Mind Lab 开源了一个长文本强化学习(RL)项目,支持 2M tokens 的长上下文。相比以往需要数千个 GPU 的 1M 上下文研究,该项目仅需 8 个 GPU 即可完成,大幅降低了超长上下文研究的算力门槛。

twitter关注列表2026-07-17#开源#技术突破#模型
观察
03

Kimi K3非对称竞争分析

Kimi K3在Arena前端/WebDev人类盲测中以1679 Elo排名第一,领先Fable 5(1631)和GPT-5.6 Sol(1618),但在综合智能指数中仅排第四(57.1分),落后Fable 5(59.9)和GPT-5.6 Sol(58.9)。K3定价15美元/百万输出tokens,远低于Fable 5的50美元,并计划7...

twitter关注列表2026-07-18#AI#模型#技术突破
观察
04

二年级学生Jo Nagai发现蝴蝶可遗传记忆

日本东京都二年级学生Jo Nagai注意到养育的食蚜 caterpillars 在蝴蝶化后仍保持对薰衣草的回避行为,经Georgetown大学Entomologist Dr. Martha Weiss合作完成实验:70%训练过的蝴蝶及其后代均表现出对薰衣草的遗传性回避,记忆在全变态发育中保存并遗传。

twitter关注列表2026-07-18#研究#技术突破#信息
值得跟进
05

AI 代码生成大势所趋

Greg Isenberg 发文指出,相比一年前的手写代码实践,如今大多数工程代码已由 AI 生成,标志着编程范式的根本转变。他援引了 Google 75% 新代码由 AI 生成、Anthropic 90%+ 代码由 Claude 编写、GitClear 代码重复率上升 81% 复用率下降 70% 等具体数据,并引用 Dario Amod...

twitter关注列表2026-07-18#技术突破#行业动态#分析
观察