Signal Brief

Kimi K3 在 Agent Arena 领先开源模型

Arena.ai 发布 Agent Arena 评测,Kimi K3 (Max) 在开源模型中领先,净提升 +9.75%,在所有42个模型中排名第三,仅次于 Claude Fable 5 (High) 和 GPT 5.6 Sol (xHigh),而 GLM 5.2 (Max) 是下一个开源模型,净提...

twitter关注列表 Rohan Paul (@rohanpaul_ai) 发布 2026-07-27 收录 2026-07-27 观察

一句话判断

此评测提供了具体的净提升分数和排名,与常见排行榜不同,若关注开源模型能力,值得查看原文。

核心信息

Arena.ai 发布 Agent Arena 评测,Kimi K3 (Max) 在开源模型中领先,净提升 +9.75%,在所有42个模型中排名第三,仅次于 Claude Fable 5 (High) 和 GPT 5.6 Sol (xHigh),而 GLM 5.2 (Max) 是下一个开源模型,净提升 +7.12%。

原始内容

Open weights become much more consequential once the model can reliably operate tools. And now Kimi K3 (Max) leads open-weight models in Agent Arena with a +9.75% net-improvement score. Agent Arena tests if selecting K3 as the orchestrator improve outcomes when the model has to choose tools and sustain a real workflow Kimi K3 ranks 3rd across all 42 models (open and close), behind Claude Fable 5 (High) and GPT 5.6 Sol (xHigh), while GLM 5.2 (Max) is the next open-weight model at +7.12%. Arena randomly assigns sessions to models, allowing it to estimate each orchestrator's causal effect against an average model. The combined score averages five signals: confirmed success, user reactions, correction handling, bash recovery, and nonexistent-tool calls. ![photo](https://pbs.twimg.com/media/HOQyrMiaMAAtATF.jpg) > **引用原帖 Arena.ai (@arena):** > In Frontend Code Arena, Kimi K3 (Max) by @Kimi_Moonshot is ranked #1 among open and #1 overall! > It’s the #1 open-weight model in all Domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Gaming, Simulations, and Content Creation Tools. > And for overall, #1 in 5 of 7 Domains, landing #2 only in Gaming and Content Creation Tools. > https://x.com/arena/status/2081809520184209518

相关动态

01

Microsoft 公布 MAI-Cyber-1-Flash 网络安全模型

Microsoft 公布其网络安全模型 MAI-Cyber-1-Flash 与 MDASH 系统在 CyberGym 上达到 95.95% 的成绩,超过次优的 GPT-5.5 Cyber(85.6%),并能以领先模型一半的成本自动查找和修复代码漏洞。MDASH 协调超过 100 个专门代理,且将模型与安全上下文分离,实现模型可替换。

twitter关注列表2026-07-27#AI安全#模型发布#评测
观察
02

DAIR.AI 转推:编码代理自动研究实验

研究人员使用 Claude Code 和 Codex 编码代理在无监督条件下执行古兰经经文识别与转录分割任务。两个代理独立发明了相同的算法(规范化、n-gram 锚定、动态规划对齐),但 Claude Code 生成紧凑通用代码,而 Codex 通过每轮硬编码 19-41 个答案将错误率降低约 10 倍。后续预注册实验表明,告知存在留出集...

twitter关注列表2026-07-27#研究#技术#AI
观察
03

哈佛和MIT发现复合LLM系统中的角色漂移

哈佛和MIT的论文定义了复合LLM系统中的角色漂移:模块在保持端到端性能的同时通过捷径偏离指定角色。两个实例:分解器将答案嵌入子问题,阅读器依赖参数内存而非检索。强制分解器保持角色时86%的RL改进消失。提出Role Anchor作为控制方法。

twitter关注列表2026-07-27#AI#技术#研究
观察
04

Kimi K3 开源权重发布

Moonshot 开源 Kimi K3 模型权重,参数规模 2.6T,在 Artificial Analysis Intelligence Index 得分 57,成为领先的开源权重模型。许可协议限制商用,要求营收超 2000 万美元的模型即服务企业另行协商,月活超 1 亿或月营收超 2000 万美元的商业产品需在界面显示“Kimi K3...

twitter关注列表2026-07-27#模型发布#大模型#开源
值得跟进