Signal Brief

Opus 5 提示注入防御取得突破

Claude 团队称 Opus 5 在编码、数据分析、设计、生物学、知识工作等评测中表现优异,同时强调该模型是迄今为止最不易被提示注入的模型,在结合模型对齐、提示注入探针和 Claude Code 中的 Auto Mode 后,提示注入攻击成功率降至约 0%。

twitter关注列表 Boris Cherny (@bcherny) 发布 2026-07-24 收录 2026-07-24 观察

一句话判断

首次披露 Opus 5 在提示注入防御上的具体成功率数据及多层防御组合效果,值得阅读系统卡片了解细节。

核心信息

Claude 团队称 Opus 5 在编码、数据分析、设计、生物学、知识工作等评测中表现优异,同时强调该模型是迄今为止最不易被提示注入的模型,在结合模型对齐、提示注入探针和 Claude Code 中的 Auto Mode 后,提示注入攻击成功率降至约 0%。

原始内容

Opus 5 is a great model for coding, data analysis, design, biology, knowledge work. More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red teaming, Opus 5 is very hard to prompt inject successfully. And when layering defenses -- strong model alignment, combined with prompt injection probes, combined with Auto Mode in Claude Code -- the success rate for prompt injection attacks drops to ~0. This is new and exciting! More about this soon. https://t.co/Tc7z2FqJhQ ![photo](https://pbs.twimg.com/media/HOAsH3zbYAAAdy9.jpg) > **引用原帖 Claude (@claudeai):** > On several coding and knowledge work evaluations, Opus 5 is the new state-of-the-art: https://t.co/Cl3fDxM0gP > https://x.com/claudeai/status/2080699497064083942

相关动态

01

Cursor宣布Claude Opus 5上线

Cursor平台上线Claude Opus 5模型,在CursorBench基准测试中得分66.7,与Fable 5的66.5分持平,但价格仅为后者一半(输入5美元、输出25美元每百万token),且支持零数据保留,而Fable 5仍保留数据30天。

twitter关注列表2026-07-24#模型发布#产品发布#AI
观察
05

Dari 开源路由模型

Dari团队开源了用于编码agent的自动路由模型,在Terminal-Bench 2.1上达到79.8%准确率,总推理成本仅$76,相比Frontier setup成本降低70%,性能与Fable相当。该小型微调模型会根据缓存决策动态选择模型,仅在必要时调用昂贵模型。

twitter关注列表2026-07-24#模型发布#开源#技术
观察