Signal Brief

Kimi K3 编码代理评测

Artificial Analysis发布Kimi K3在编码代理指数上的评测结果:得分57,排名第5,性能与GPT-5.6 Terra和GPT-5.5持平(57),超过Opus 4.8(55),略低于Grok 4.5(58)和Fable 5(59);成本平均3.18美元/任务,比GPT-5.6 S...

twitter关注列表 Artificial Analysis (@ArtificialAnlys) 发布 2026-07-17 收录 2026-07-17 观察

一句话判断

该评测提供了详细的成本对比和性能细分,值得点开原文查看完整指数及方法。

核心信息

Artificial Analysis发布Kimi K3在编码代理指数上的评测结果:得分57,排名第5,性能与GPT-5.6 Terra和GPT-5.5持平(57),超过Opus 4.8(55),略低于Grok 4.5(58)和Fable 5(59);成本平均3.18美元/任务,比GPT-5.6 Sol便宜55%。

原始内容

Kimi K3 in Kimi Code CLI scores 57 and ranks #5 on the Artificial Analysis Coding Agent Index. Its performance is just behind Grok 4.5, in line with GPT-5.6 Terra and GPT-5.5, and ahead of Opus 4.8 Key results: ➤ Joint #5 overall on the Artificial Analysis Coding Agent Index: Kimi K3 scores 57, matching GPT-5.6 Terra max (57) and GPT-5.5 xhigh (57). It outperforms Opus 4.8 max (55), sits just behind Grok 4.5 high (58) and Fable 5 max (59), and trails GPT-5.6 Sol max (61). ➤ Strong performance across all three coding evaluations: K3 scores 84% on Terminal-Bench v2, 64% on DeepSWE, and 23% on SWE-Atlas-QnA. It outperforms Fable 5 on Terminal-Bench, Grok 4.5 and Opus 4.8 on DeepSWE, and GPT-5.6 Terra and GPT-5.5 on SWE-Atlas-QnA. ➤ Cost-efficient frontier coding performance: K3 costs an average of $3.18 per task. It is 55% cheaper than GPT-5.6 Sol max ($7.08), 73% cheaper than Fable 5 max ($11.72), 37% cheaper than GPT-5.5 xhigh ($5.07), and 59% cheaper than Opus 4.8 max ($7.70). Grok 4.5 high ($2.59) and GPT-5.6 Terra max ($2.76) are slightly cheaper. ➤ Leads open weight coding models (assuming weights are released): K3’s score of 57 is substantially ahead of other tested open-weight configurations, including GLM-5.2 at 40 and DeepSeek V4 Pro at 29. ![photo](https://pbs.twimg.com/media/HNdbIz9bAAAkOnJ.jpg) Artificial Analysis (@ArtificialAnlys): Further coding agent benchmarks on Artificial Analysis: https://t.co/huXZWndXsZ

相关动态

01

Kimi K3 在 SpreadsheetBench 2 上排名第一

Kimi K3 在 SpreadsheetBench 2 基准测试中排名第一,超越 Claude Fable 5,完成 34.8% 的 workflow 任务。该基准测试涉及 321 个专家策划任务,平均每个任务包含 11.8 个工作表和 593.5 个单元格更改,聚焦于完整的电子表格工作簿执行。

twitter关注列表2026-07-18#模型#技术突破#评测
观察
02

Runway Agent 在第三方评测中全面领先

Physion Labs 发布 Physion-Arc 1.0 基准测试,对 Runway、Luma、MiniMax、Kling、Utopia 和 TapNow 六款 AI 视频代理进行独立人类评估,使用 30 个电影提示和 16 个指标。Runway Agent 2.0 在叙事连贯性、电影语言和制作质量三项核心维度全部排名第一。

twitter关注列表2026-07-17#评测#多模态#模型
观察
03

Grok 4.5成本效率对比评测

Artificial Analysis数据显示,Grok 4.5每次任务成本仅0.31美元,而Claude Fable 5为2.75美元、Claude Opus 4.8为1.80美元、GPT-5.6 Sol为1.04美元、Kimi K3为0.95美元,Grok 4.5成本分别低约9倍、6倍、3倍和3倍;在FrontierSWE上,Grok...

twitter关注列表2026-07-17#AI#模型#评测
观察
04

8天内四款前沿模型发布评测

Artificial Analysis 发布评测,8天内 SpaceXAI 的 Grok 4.5(54分)、OpenAI 的 GPT-5.6 Sol/Terra/Luna(最高59/55/51)、Meta 的 Muse Spark 1.1(51分)和 Moonshot AI 的 Kimi K3(57分)相继发布,前沿模型实验室从6月初的2...

twitter关注列表2026-07-17#模型发布#评测#技术突破
值得跟进
05

GMI Cloud 测试 Kimi K3 vs Claude Fable 5 3D 建模

GMI Cloud 对 Kimi K3 和 Claude Fable 5 进行 3D 建模对比测试:像素风格下 Kimi K3 耗时 1h11min、花费 $14.5,Fable 5 耗时 49min、花费 $21;原始风格下 Kimi K3 耗时 1h17min、花费 $19.3,Fable 5 耗时 1h5min、花费 $47.2。K...

twitter关注列表2026-07-16#模型#评测#技术
观察