Signal Brief

Kimi K3 表现亮眼及 AI 领域多项动态

Kimi K3 在自治法律工作基准上几乎超越 Claude Fable 5 两倍,并修复了 15 个 Codex 和 Fable 拒绝修复的严重安全漏洞。开放权重模型在长周期网络能力上落后闭源前沿的时间从 6-10 个月缩短至 4-7 个月。American companies 能以中国竞争对手十分...

twitter关注列表 Rohan Paul (@rohanpaul_ai) 发布 2026-07-21 收录 2026-07-21 观察

一句话判断

包含 Kimi K3 与 Claude 的量化对比及开放权重模型追赶进度的具体数据,值得点开原文查看详细基准和成本分析。

核心信息

Kimi K3 在自治法律工作基准上几乎超越 Claude Fable 5 两倍,并修复了 15 个 Codex 和 Fable 拒绝修复的严重安全漏洞。开放权重模型在长周期网络能力上落后闭源前沿的时间从 6-10 个月缩短至 4-7 个月。American companies 能以中国竞争对手十分之一的成本部署 Kimi K3,但 K3 因需求过重而暂停新订阅。

原始内容

Today’s edition of my newsletter just went out. 🔗 https://t.co/Eo60JWRzFs 🗞️ On long-horizon cyber capability, leading open-weight models now trail the closed frontier by only 4 to 7 months, down from 6 to 10 months through much of 2025. 🗞️ “AI advice suppresses people’s willingness to say “I don’t know”, even when the advice is wrong and accuracy is incentivized” 🗞️ Kimi K3 nearly doubled its nearest rival Claude Fable 5, on a demanding benchmark for autonomous legal work. 🗞️ Kimi K3 just fixed 15 critical security bugs that Codex and Fable refused because of “cyber guardrails.” 🗞️ “AI Agents Do Not Fail Alone:The Context Fails First” 🗞️ “American companies such as Modal, Fireworks, and Baseten will be able to serve Kimi K3, at one-tenth the cost of their Chinese competitors because they have access to advanced Nvidia and AMD chips. 🗞️ Kimi K3 is facing a compute crunch. Demand is too heavy, so new subscriptions are currently blocked. 🗞️ Why One OpenAI Senior Employee Thinks Open-Weight Models Are Decelerationist ![photo](https://pbs.twimg.com/media/HNtiTpCbgAAnWth.png)

相关动态

01

微软公布AI在科学发现的投入

微软与美国能源部通过#GenesisMission深化合作,投资AI基础设施、科学计算、工程专长及合作枢纽,帮助全国实验室、大学和产业合作。目标是加速AI驱动的科学发现,覆盖能源、医药和先进材料,以增强美国竞争力并创造长期价值。

twitter关注列表2026-07-22#技术#政策#创新
观察
02

Alphabet Q2 AI业绩增长

Alphabet公布Q2业绩:收入同比增长24%,Google Cloud增长82%,Gemini app月活达9.5亿,模型API处理量22B tokens/min(上季度16B+),Gemini Enterprise被90%财富100强采用。

twitter关注列表2026-07-22#大模型#行业动态#技术
观察
03

利用 Embedding 模型实现图像打标

Han Xiao 通过将冻结的 jina-v5-omni 多模态嵌入模型进行测试时缩放(scaled at test time),在无需训练、无需第二模型及外部知识的硬约束下,实现了强大的开放词汇多标签 n-gram 图像打标器。

twitter关注列表2026-07-22#技术#研究#模型
观察