Signal Brief

四款AI模型IMO 2026获满分42/42

Claude Fable 5、GPT 5.6 Sol、Kimi K3和Axiom Math四款AI模型在IMO 2026中获得满分42/42。Fable 5用1次尝试最快,Sol用1次尝试最便宜,K3用4次尝试消耗大量token,Axiom Math在Lean中证明了所有题目。学生需9小时完成6题,...

twitter关注列表 Deedy (@deedydas) 发布 2026-07-21 收录 2026-07-21 值得跟进

一句话判断

新增了多个前沿模型在IMO上的具体表现对比,包括Fable速度、Sol成本、Axiom的Lean证明能力,值得关注原文的详细数据和推理过程。

核心信息

Claude Fable 5、GPT 5.6 Sol、Kimi K3和Axiom Math四款AI模型在IMO 2026中获得满分42/42。Fable 5用1次尝试最快,Sol用1次尝试最便宜,K3用4次尝试消耗大量token,Axiom Math在Lean中证明了所有题目。学生需9小时完成6题,Fable和Sol用时不到4小时。

原始内容

The International Math Olympiad (IMO) 2026, the hardest math contest for high schoolers, just ended. I ran Fable (high), Sol (xhigh), K3 (max) and Axiom against it and all got a perfect score of 42/42 (repo below if you want to check their solutions): — Claude Fable 5 was the solved it in 1 attempt, and was the fastest. — GPT 5.6 Sol took 1 more attempts, and was cheapest. — Kimi K3 did it but took 4 more attempts, and took a LOT of tokens. — Axiom Math actually proved everything in Lean. P3 and P6 were the hardest followed by P2, judging by attempts + num tokens. Students had 9hrs to solve these 6 problems, and Fable and Sol were under 4hrs. The frontier of AI has officially moved well past IMO math. ![photo](https://pbs.twimg.com/media/HNuL4q3aQAIMcXc.jpg)

相关动态

01

NVIDIA Vera Rubin 平台发布

NVIDIA 发布 Vera Rubin 平台,性能功耗比提升 10 倍;CoreWeave 等云服务商部署 Vera Rubin NVL72,每兆瓦 token 数比 Blackwell 多 10 倍;DeepInfra 基准测试显示 Vera CPU 速度是其他 CPU 的 2 倍以上,支持更多并发 AI 代理。

twitter关注列表2026-07-21#技术突破#产品发布#AI
值得跟进
04

Grok 4.5 集成 Microsoft Outlook

Grok 4.5 现可直接在 Microsoft Outlook 中使用,支持总结邮件线程与附件、识别决策与待办任务、以用户语气起草回复、搜索网络和 𝕏,以及整理、归档、删除或标记邮件。

twitter关注列表2026-07-21#AI#产品更新#大模型
观察
05

Anthropic 1.5亿美元和解作者版权诉讼

美国法官批准Anthropic支付1.5亿美元给作者,作为其使用盗版书籍训练Claude的版权和解金。这是美国版权案件中最大和解,涉及700万本盗版书,91%作者已领取份额。法院裁定训练属于合理使用,但存储盗版副本构成侵权。

twitter关注列表2026-07-21#AI#行业动态#政策
观察