Signal Brief

Microsoft 公布 MAI-Cyber-1-Flash 网络安全模型

Microsoft 公布其网络安全模型 MAI-Cyber-1-Flash 与 MDASH 系统在 CyberGym 上达到 95.95% 的成绩,超过次优的 GPT-5.5 Cyber(85.6%),并能以领先模型一半的成本自动查找和修复代码漏洞。MDASH 协调超过 100 个专门代理,且将模型...

twitter关注列表 Rohan Paul (@rohanpaul_ai) 发布 2026-07-27 收录 2026-07-27 观察

一句话判断

差异点在于微软将模型与安全上下文分离,实现模型可替换性,并整合超过100个专业代理协作,这比单纯模型性能提升更具架构创新意义。

核心信息

Microsoft 公布其网络安全模型 MAI-Cyber-1-Flash 与 MDASH 系统在 CyberGym 上达到 95.95% 的成绩,超过次优的 GPT-5.5 Cyber(85.6%),并能以领先模型一半的成本自动查找和修复代码漏洞。MDASH 协调超过 100 个专门代理,且将模型与安全上下文分离,实现模型可替换。

原始内容

Microsoft is reporting 95.95% on CyberGym for its MDASH configuration. CyberGym measures whether AI agents can reproduce real software vulnerabilities from code. The next-best result shown is GPT-5.5 Cyber at 85.6%. Gemini 3.5, GPT-5.6 Sol, and Mythos 5 all sit around 83% to 84%. MAI-Cyber-1-Flash is the AI model. MDASH is the larger agent-and-orchestration system that uses it. It coordinates more than 100 specialised agents and gives them different roles, tools, prompts and stopping rules. If it really works, it can automate the slow, expensive work of finding and fixing hidden flaws in huge codebases. The harness, security context, signals, and action space are kept separate from the model family. In principle, that makes the underlying model replaceable. Microsoft can introduce a cybersecurity-specific model, combine it with another model where needed, and keep the surrounding investigation and remediation workflow intact. One particular sentence in their official blog is particularly interesting. "Microsoft sees more than 100 trillion security signals every day" ![photo](https://pbs.twimg.com/media/HOQsj0dbMAA5R8I.png) > **引用原帖 Satya Nadella (@satyanadella):** > Today, we are announcing a series of updates that give customers frontier-grade security at half the cost. > MAI-Cyber-1-Flash is our first cybersecurity model, built ground up to find the most challenging vulnerabilities in complex code bases. When combined with MDASH, it delivers world-class performance at 50 percent of the cost of leading models. > We are bringing this capability to market through Project Perception, a complete agentic security offering grounded in real-world signals and security workflows. Teams of specialized agents work together to simulate attacks, detect and triage/investigate, and fix and remediate. > This is the benefit of building the harness, context/signals, and action space separate from one model family. By combining specialized models and data with the right agents, tools, security context, and harness, we can advance the frontier of cost to outcome. > https://x.com/satyanadella/status/2081779755146482153

相关动态

01

Kimi K3 在 Agent Arena 领先开源模型

Arena.ai 发布 Agent Arena 评测,Kimi K3 (Max) 在开源模型中领先,净提升 +9.75%,在所有42个模型中排名第三,仅次于 Claude Fable 5 (High) 和 GPT 5.6 Sol (xHigh),而 GLM 5.2 (Max) 是下一个开源模型,净提升 +7.12%。

twitter关注列表2026-07-27#评测#开源#技术
观察
03

Kimi K3技术报告发布

Kimi.ai发布Kimi K3技术报告,该模型为2.8T参数MoE,原生支持视觉理解,拥有百万token上下文窗口。报告详细描述了新架构Kimi Delta Attention,通过移除位置编码、注意力残差、分层训练等技术,实现约2.5倍于Kimi K2的扩展效率。

twitter关注列表2026-07-27#模型发布#技术突破#大模型
值得跟进
04

Kimi K3 开源权重发布

Moonshot 开源 Kimi K3 模型权重,参数规模 2.6T,在 Artificial Analysis Intelligence Index 得分 57,成为领先的开源权重模型。许可协议限制商用,要求营收超 2000 万美元的模型即服务企业另行协商,月活超 1 亿或月营收超 2000 万美元的商业产品需在界面显示“Kimi K3...

twitter关注列表2026-07-27#模型发布#大模型#开源
值得跟进