Signal Brief

英美政府安全机构联合评估Kimi K3网络安全能力

英美政府安全机构联合发布报告,对比Kimi K3与美国前沿模型的网络安全能力。Kimi K3在模拟攻击中平均步骤17,而美国最强模型达28.5;在ExploitBench中,Kimi K3评分32.2%,美国领先模型76.2%,GLM-5.2为24.4%。

twitter关注列表 Rohan Paul (@rohanpaul_ai) 发布 2026-07-23 收录 2026-07-23 观察

一句话判断

报告提供了Kimi K3与美国前沿模型在网络安全能力上的量化差距,以及开源模型的对比数据,建议关注具体评估方法。

核心信息

英美政府安全机构联合发布报告,对比Kimi K3与美国前沿模型的网络安全能力。Kimi K3在模拟攻击中平均步骤17,而美国最强模型达28.5;在ExploitBench中,Kimi K3评分32.2%,美国领先模型76.2%,GLM-5.2为24.4%。

原始内容

British and American government safety institutes jointly published a report comparing Kimi K3 versus top frontier US models. Kimi K3 stopped at step 17 on average, while the strongest American models reached 28.5. Kimi K3 remains substantially behind frontier U.S. models in offensive cyber capability, but it is stronger than the previous leading open-weight model. Can autonomously execute meaningful portions of an attack, occasionally completes an entire simulated enterprise attack, and does not reliably refuse offensive requests. On the ExploitBench evaluation: Leading U.S. models scored 76.2%. Kimi K3 scored 32.2%. GLM-5.2 scored 24.4%. American closed models were measured with safeguards switched off, so those numbers show ceilings, not shipping products. ![photo](https://pbs.twimg.com/media/HN78q0fakAAwiUq.png) ![photo](https://pbs.twimg.com/media/HN78tzlbIAASrqc.png) ![photo](https://pbs.twimg.com/media/HN78wJTbsAAd2qB.png) > **引用原帖 U.S. Department of Commerce (@CommerceGov):** > CAISI’s latest blog post evaluates Kimi K3 and its cyber capabilities. > Based on a preliminary cyber-focused evaluation, Kimi K3 performed significantly below the leading U.S. frontier AI models. > https://t.co/r9K3Pp0IiH > https://x.com/CommerceGov/status/2080341953086886387

相关动态

01

OpenAI 推出桌面版 ChatGPT Voice 语音控制功能

OpenAI 在 macOS 和 Windows 桌面应用中推出 ChatGPT Voice 功能,基于 GPT-Live 模型,支持用户通过语音控制电脑、同时指挥多个 agents,实现语音交互与任务协调。该功能面向 Plus、Pro、Business、Edu 和 Enterprise 计划用户,全球逐步上线。

twitter关注列表2026-07-23#AI#产品更新#模型
观察
02

AI医疗索赔代理

PE支持的医疗账单平台在四个月内部署了七个生产级AI代理,处理真实医疗索赔并确保零PHI泄露。他们通过构建包含34个动态变量的丰富化层、固定架构和严格的输出模式验证,将调用路由在两个模型提供商之间进行交换而不影响输出。与常规直接接入模型的做法相比,该团队花费的工程时间主要用于构建评估套件和数据预处理,实现了可靠的生产运行。

twitter关注列表2026-07-23#技术#评测#发布
观察
05

红点笔记模型获 IMO 完美分数

RedNote 的 dots‑note‑3.0 AI 模型在国际数学奥林匹克赛中得满分 42/42,超越去年 DeepMind 和 OpenAI 的 35/42,且仅有 7 名 666 名人类选手匹配。该模型仍在 beta,为 dots3 系列中最小的变体,公司承诺日后开源。

twitter关注列表2026-07-23#模型#技术
值得跟进