Signal Brief

Fable 5 自动化 16.1% 远程工作,RLI 最新结果

CAIS和Scale公布Remote Labor Index最新结果,Fable 5模型自动化16.1%的真实远程工作项目,约为Opus 4.8(8.3%)的2倍,GPT-5.5为6.3%。RLI诞生时最佳模型仅2.5%,进步显著,且任务涉及CAD、架构等复杂工具使用,而非简单文本。报告指出代理设置...

twitter关注列表 Rohan Paul (@rohanpaul_ai) 发布 2026-07-05 收录 2026-07-06 观察

一句话判断

该结果展示了代理架构改进带来的巨大提升,而非单纯模型能力,值得关注RLI基准及代理设置的具体方法。

核心信息

CAIS和Scale公布Remote Labor Index最新结果,Fable 5模型自动化16.1%的真实远程工作项目,约为Opus 4.8(8.3%)的2倍,GPT-5.5为6.3%。RLI诞生时最佳模型仅2.5%,进步显著,且任务涉及CAD、架构等复杂工具使用,而非简单文本。报告指出代理设置(工具使用、桌面环境、长时运行、工作者-批评者循环)比模型本身更关键。

原始内容

CAIS and Scale (AI safety research group) say Fable 5 now automates 16.1% of real remote-work projects, about 2x Opus 4.8. Remote Labor Index tests whether an AI can finish paid freelance work well enough for a client to accept it. Each task comes with briefs, files, and a professional deliverable used as the human baseline. Fable 5 led the new results, while Opus 4.8 reached 8.3% and GPT-5.5 reached 6.3%. That 16.1% number still means failure on most tasks, but the direction of improvement is massive. For the context of how large this jump is, the best model scored only 2.5% when RLI launched. The work is not toy prompting, since tasks include CAD, architecture, animation, audio, data analysis, and web apps. This means the benchmark is testing computer work across messy tools, not narrow text answers. Fable 5 looked strongest on examples like ring modeling, animation, and bathroom design. The automated judge ranked models well, but overstated GPT-5.5 by about 2.9x and Opus 4.8 by about 2.3x. The result says AI agents are improving fast, but quality control remains the hard wall. Another key point is that the model alone is not the whole product anymore. The gains are coming from stronger agent setups: better tool use, full desktop environments, professional software, longer runtimes, and worker-critic loops where 1 agent does the task and another reviews it like a demanding client. ![photo](https://pbs.twimg.com/media/HMfepFDb0AAbBk-.png) > **引用原帖 Rohan Paul (@rohanpaul_ai):** > AI revenue is scaling 3x quicker than mobile or internet did. > The unit of value is shifting from attention to completed work. https://t.co/WrOhHzyzow > https://x.com/rohanpaul_ai/status/2073809550403408171

相关动态

02

Kimi K3 在 SpreadsheetBench 2 上排名第一

Kimi K3 在 SpreadsheetBench 2 基准测试中排名第一,超越 Claude Fable 5,完成 34.8% 的 workflow 任务。该基准测试涉及 321 个专家策划任务,平均每个任务包含 11.8 个工作表和 593.5 个单元格更改,聚焦于完整的电子表格工作簿执行。

twitter关注列表2026-07-18#模型#技术突破#评测
观察
03

Mind Lab 开源长文本强化学习项目

Mind Lab 开源了一个长文本强化学习(RL)项目,支持 2M tokens 的长上下文。相比以往需要数千个 GPU 的 1M 上下文研究,该项目仅需 8 个 GPU 即可完成,大幅降低了超长上下文研究的算力门槛。

twitter关注列表2026-07-17#开源#技术突破#模型
观察
04

Kimi K3非对称竞争分析

Kimi K3在Arena前端/WebDev人类盲测中以1679 Elo排名第一,领先Fable 5(1631)和GPT-5.6 Sol(1618),但在综合智能指数中仅排第四(57.1分),落后Fable 5(59.9)和GPT-5.6 Sol(58.9)。K3定价15美元/百万输出tokens,远低于Fable 5的50美元,并计划7...

twitter关注列表2026-07-18#AI#模型#技术突破
观察
05

二年级学生Jo Nagai发现蝴蝶可遗传记忆

日本东京都二年级学生Jo Nagai注意到养育的食蚜 caterpillars 在蝴蝶化后仍保持对薰衣草的回避行为,经Georgetown大学Entomologist Dr. Martha Weiss合作完成实验:70%训练过的蝴蝶及其后代均表现出对薰衣草的遗传性回避,记忆在全变态发育中保存并遗传。

twitter关注列表2026-07-18#研究#技术突破#信息
值得跟进