Signal Brief

Fugu-Ultra v1.1发布

Fugu发布Fugu-Ultra v1.1,性能最高提升7.9分,尤其在ProgramBench和Terminal Bench 2.1上,价格不变。

twitter关注列表 Sakana AI (@SakanaAILabs) 发布 2026-07-24 收录 2026-07-24 观察

一句话判断

新增了具体基准提升细节,值得查看原文。

核心信息

Fugu发布Fugu-Ultra v1.1,性能最高提升7.9分,尤其在ProgramBench和Terminal Bench 2.1上,价格不变。

原始内容

Announcing Fugu-Ultra v1.1 🐡 We’ve been thrilled by the reception to the Fugu model family. Thanks to everyone who tried it, shared feedback, and trusted Fugu with real work. Today, we’re releasing Fugu-Ultra v1.1 → https://t.co/hhO6qTawgb Upgraded to incorporate the latest frontier models, resulting in stronger performance across every benchmark shown, including gains of up to 7.9 points over v1.0, with particularly strong results on ProgramBench and Terminal Bench 2.1. Fugu-Ultra v1.1 is more capable across coding, agentic tasks, and advanced reasoning, and available at the same price as Fugu-Ultra v1.0 The frontier keeps moving, and Fugu keeps getting better. ![photo](https://pbs.twimg.com/media/HN89IPDaEAEGy41.jpg)

相关动态

02

When Do Multi-Agent Systems Help? An Information Bottleneck Perspective

该论文通过信息瓶颈视角,在5个基准测试、3种模型大小、18个测试中,比较了单agent、单agent遵循任务拆分、多agent系统的性能。发现多agent系统在信息传递紧凑完整时有效,尤其对弱模型,但当后续步骤需要精确早期细节时优势减弱或逆转。

twitter关注列表2026-07-24#AI#技术#分析
观察
05

BestBlogs 早报 · 07-24

本次早报包含10条AI新闻:Anthropic升级Claude语音模式支持Opus与Sonnet;梁文锋透露DeepSeek将Agent持续学习作为下一目标,承认算力与数据约束;阿里开源Agent评测框架skill-up,内部将1200行代码收敛为声明式;Andrew Ng推出开源智能体OpenWorker;此外还有关于Agent PR频...

twitter关注列表2026-07-23#AI#行业动态#技术更新
观察