Signal Brief

Anthropic前员工谈开放模型与安全风险

前Anthropic员工Noah Lebovic改变观点,认为开放模型已足够强大,例如Opus 4.6和GLM 5.1/5.2在渗透测试中表现优于Anthropic模型。他观察到恶意行为者仍使用Claude Code或Codex订阅,而非开放模型,并爆料Anthropic的GTM人员通过大额合同降低...

twitter关注列表 Thomas Wolf (@Thom_Wolf) 发布 2026-07-27 收录 2026-07-27 值得跟进

一句话判断

值得阅读原文,获取内部视角对AI安全风险与开放模型实际使用的独特见解。

核心信息

前Anthropic员工Noah Lebovic改变观点,认为开放模型已足够强大,例如Opus 4.6和GLM 5.1/5.2在渗透测试中表现优于Anthropic模型。他观察到恶意行为者仍使用Claude Code或Codex订阅,而非开放模型,并爆料Anthropic的GTM人员通过大额合同降低安全限制,导致开放模型更可靠且不易受限制。

原始内容

👀 > **引用原帖 Noah Lebovic (@NoahLebovic):** > I don't think it's a lack of imagination. I also used to work at Anthropic, think trends will continue, and used to agree with this. But I've changed my mind and now disagree with this take. > Open models are already capable enough to do what you described. For example, I used Opus 4.6 to gain access to other folks medical records, hijack bank accounts, etc. back in February. GLM 5.1 is more capable than Opus 4.6 in most pentesting environments, and it came out in April. > Despite capable open-weight models existing, the sketchier folks I know are still using a Claude Code or Codex subscription for hacking. (Even well-resourced groups in other countries! They use the grey/black market of discounted Ant/OAI subscription tokens sold through resellers.) So I see most of the materialized risk here as still coming from Anthropic and OpenAI; safeguards aren't sufficient to stop a moderately dedicated actor. > The groups I know who are using open-weight models are legitimate offensive security companies. They won't break the rules to use subscription-based pricing, the open-weight models are more reliable in that they don't require specific jailbreaks nor hit classifiers, and the labs use massive partnerships or spend as a prereq for lowering classifiers/safeguards. I know of three legitimate groups running GLM 5.2 as their primary model. > That last part applies for Anthropic, too: I know of two instances where two different Anthropic GTM people used large comitted spend contracts as a prereq for lowering safeguards, and I directly witnessed one. > On the inside, I know the narrative and intent is genuinely about safety. But from the outside, Anthropic-the-system seems to be optimizing for revenue and control/power, isn't diffusing capabilities to defenders, and also doesn't have adequate safeguards to prevent misuse from dedicated bad actors. > As a result, I now lean towards a future where capable open models are freely available (at least for cyber, bio is harder); I don't trust Anthropic or other frontier labs to handle this sufficiently well without diffused capabilties given what I've seen so far. > https://x.com/NoahLebovic/status/2081277517709922501

相关动态

01

From Human-Centric to Agentic Code Review: The Impact of Different Generations of Generative AI Technology on Review Quality

GitHub 上 207 个项目的 102 万条拉取请求研究显示,逐步或快速采用 AI 代理可将审查时间缩短 2.5 至 4.5 天/KLOC,而早期大量使用 LLM 审阅未带来显著效率提升,且 AI 审阅导致审阅气味在 78% 至 94% 的请求中出现,高于人工审阅的 69% 至 76%。

twitter关注列表2026-07-27#技术突破#行业动态
值得跟进
02

Human Capital, AI, and Labor Commoditization

一项研究分析AI对自由职业市场的影响,发现在ChatGPT后,AI暴露岗位中人力资本信号重要性下降7.8%,价格信号重要性上升1.1%,导致需求向更廉价工人转移,支持AI使这些工人更可替代的观点。

twitter关注列表2026-07-26#AI#研究#行业动态
观察
03

DeepSeek 暂停第二轮融资

DeepSeek reportedly paused its second fundraising round targeting a valuation near $74B, after raising about $7.4B in June. The pause is partly due to Liang Wenfeng's fru...

twitter关注列表2026-07-26#行业动态#大模型
观察
04

中国监管AI伴侣

中国出台新规限制AI伴侣服务,要求提供商遏制操控性依恋、识别过度依赖并提醒用户对话对象为AI;数据表明每日情感聊天减少对人类支持的10.3%;同时禁止向未成年人提供虚拟伴侣或亲属。

twitter关注列表2026-07-26#政策#AI安全#行业动态
观察