Signal Brief

Offloop多智能体超越Claude Code

Offloop 4 人团队公布其多智能体 harness 在 GDPval 基准上以 84.9 分、单任务成本 $1.65 超越 Claude Code(Opus 4.8 得 82.4 分、$14.38) 与 Codex(GPT 5.6 Sol 得 83.3 分、$5.20),并声称在 GDP.pd...

twitter关注列表 Rohan Paul (@rohanpaul_ai) 发布 2026-07-23 收录 2026-07-24 观察

一句话判断

首个在 GDPval 上以大幅成本优势击败 Claude Code 与 Codex 的多智能体系统,值得关注其架构细节与复现可能性。

核心信息

Offloop 4 人团队公布其多智能体 harness 在 GDPval 基准上以 84.9 分、单任务成本 $1.65 超越 Claude Code(Opus 4.8 得 82.4 分、$14.38) 与 Codex(GPT 5.6 Sol 得 83.3 分、$5.20),并声称在 GDP.pdf(44 分) 与 JobBench(67.2 分) 亦达 SOTA,GDPval 覆盖 44 个职业、9 大行业、对应约 $2.4 万亿美元美国知识工作薪资。

原始内容

So much recent work and research papers points to the same thing: the "harness" is becoming the real capability layer. @Offloop 's 4-person team demonstrated a multi-agent harness outperforming Claude Code and Codex on GDPval benchmarks, targeting $2.4T in US knowledge work. - The team scored 84.9 at $1.65 per task. - Opus 4.8 inside Claude Code scored 82.4, and GPT 5.6 Sol inside Codex scored 83.3, costing far more per task, $14.38 and $5.20 GDPval measures how well AI handles real work across 44 occupations and 9 major industries. Models get shell access and web browsing, then face blind pairwise comparisons against human experts. And those tasks map onto US jobs paying roughly $2.4T a year. ![photo](https://pbs.twimg.com/media/HN7gmVbbYAAEFSP.jpg) > **引用原帖 Offloop (@Offloop):** > In our internal evaluations, Offloop achieved new state of the art results on three benchmarks evaluating professional knowledge work: > 84.9 on GDPval, 44 on GDP.pdf, 67.2 on JobBench - each ahead of Codex and Claude Code, at 1/3 to 1/10 of their cost per task. > Lower cost per completed task means agents you can leave running, more tasks per dollar, and margins that survive scale. > https://x.com/Offloop/status/2080304786646433874

相关动态

03

SANA-Video 2.0 发布

Enze Xie 团队发布 SANA-Video 2.0,采用混合线性-softmax 注意力、块注意力残差和 Sol-Engine 加速,统一 5B 和 14B 模型,在 16 节点 H100 上训练 5B 模型,VBench 总分 84.30,单 H100 生成 720p/5s 仅需 13.06s,速度比 Wan 2.2-A14B 快...

twitter关注列表2026-07-24#模型发布#技术突破#多模态
观察
05

Ant Ling 发布 Ling-3.0-flash 模型

Ant Ling 发布 Ling-3.0-flash 混合推理 MoE 模型,124B 参数,仅 5.1B 活跃参数,采用 KDA+MLA 混合注意力机制,支持 256K 上下文。以 1/8 总参数量和 1/12 活跃参数量,在多数基准上匹配或超越其 1T 旗舰模型。SGLang 正与其团队合作提供 day-0 支持。

twitter关注列表2026-07-23#模型发布#大模型#技术突破
观察