Signal Brief

开源模型长期网络能力差距缩小至4-7个月

长期网络能力方面,领先开源模型落后闭源前沿从2025年大部分时间的6-10个月缩短至4-7个月。GLM-5.2匹配了约4个月前发布的闭源模型。AI安全研究所的32步“The Last Ones”任务中,GPT-5.6 Sol在10次尝试中完成7次,每次预算1亿token,且性能随推理token增加而...

twitter关注列表 Rohan Paul (@rohanpaul_ai) 发布 2026-07-17 收录 2026-07-18 观察

一句话判断

与先前认知相比,开源模型追赶速度加快,且长期网络能力表现出可随推理计算扩展的特性,值得关注GLM-5.2的具体评测及GPT-5.6 Sol的防御应用。

核心信息

长期网络能力方面,领先开源模型落后闭源前沿从2025年大部分时间的6-10个月缩短至4-7个月。GLM-5.2匹配了约4个月前发布的闭源模型。AI安全研究所的32步“The Last Ones”任务中,GPT-5.6 Sol在10次尝试中完成7次,每次预算1亿token,且性能随推理token增加而提升。

原始内容

On long-horizon cyber capability, leading open-weight models now trail the closed frontier by only 4 to 7 months, down from 6 to 10 months through much of 2025. GLM-5.2 matched closed models released about 4 months earlier. On long-horizon cyber ranges, GLM-5.2 matched Claude Opus 4.5, released roughly 7 months earlier. And also that long-horizon cyber capability can now be scaled with compute. On the AI Security Institute’s 32-step “The Last Ones” range, GPT-5.6 Sol completed the full 32-step range in 7 of 10 attempts. With a 100M-token budget for each run. Performance continued improving as the model received more inference tokens. i.e. operators can gain materially stronger cyber capability simply by spending more runtime compute, without retraining the model. ![photo](https://pbs.twimg.com/media/HNdxz1MbEAA7WSS.jpg) ![photo](https://pbs.twimg.com/media/HNdzFTuaYAA8leW.png) > **引用原帖 OpenAI (@OpenAI):** > GPT-5.6 Sol sets a new state of the art in cybersecurity on “The Last Ones” cyber range. > We’re already seeing that capability translate into defensive outcomes: helping teams find, validate, and fix vulnerabilities in real-world code. > Put it to work with Codex Security: https://t.co/Fvz9wpLjrt > https://x.com/OpenAI/status/2078243667081617826

相关动态

01

Kimi K3 在 SpreadsheetBench 2 上排名第一

Kimi K3 在 SpreadsheetBench 2 基准测试中排名第一,超越 Claude Fable 5,完成 34.8% 的 workflow 任务。该基准测试涉及 321 个专家策划任务,平均每个任务包含 11.8 个工作表和 593.5 个单元格更改,聚焦于完整的电子表格工作簿执行。

twitter关注列表2026-07-18#模型#技术突破#评测
观察
02

Mind Lab 开源长文本强化学习项目

Mind Lab 开源了一个长文本强化学习(RL)项目,支持 2M tokens 的长上下文。相比以往需要数千个 GPU 的 1M 上下文研究,该项目仅需 8 个 GPU 即可完成,大幅降低了超长上下文研究的算力门槛。

twitter关注列表2026-07-17#开源#技术突破#模型
观察
03

Kimi K3非对称竞争分析

Kimi K3在Arena前端/WebDev人类盲测中以1679 Elo排名第一,领先Fable 5(1631)和GPT-5.6 Sol(1618),但在综合智能指数中仅排第四(57.1分),落后Fable 5(59.9)和GPT-5.6 Sol(58.9)。K3定价15美元/百万输出tokens,远低于Fable 5的50美元,并计划7...

twitter关注列表2026-07-18#AI#模型#技术突破
观察
04

二年级学生Jo Nagai发现蝴蝶可遗传记忆

日本东京都二年级学生Jo Nagai注意到养育的食蚜 caterpillars 在蝴蝶化后仍保持对薰衣草的回避行为,经Georgetown大学Entomologist Dr. Martha Weiss合作完成实验:70%训练过的蝴蝶及其后代均表现出对薰衣草的遗传性回避,记忆在全变态发育中保存并遗传。

twitter关注列表2026-07-18#研究#技术突破#信息
值得跟进
05

AI 代码生成大势所趋

Greg Isenberg 发文指出,相比一年前的手写代码实践,如今大多数工程代码已由 AI 生成,标志着编程范式的根本转变。他援引了 Google 75% 新代码由 AI 生成、Anthropic 90%+ 代码由 Claude 编写、GitClear 代码重复率上升 81% 复用率下降 70% 等具体数据,并引用 Dario Amod...

twitter关注列表2026-07-18#技术突破#行业动态#分析
观察