Signal Brief

GPT-5.6 Sol (max) 在 AA-Briefcase 获得最高 Presentation Elo

GPT-5.6 Sol (max) 在 AA-Briefcase 智能体知识工作基准测试中获得最高 Presentation Elo 分数,较GPT-5.5 (xhigh)提升约500分,预计在对比中胜率约95%。AA-Briefcase综合二元评级检查与分析质量及演示文稿成对比较,测试模型在复杂项...

twitter关注列表 Artificial Analysis (@ArtificialAnlys) 发布 2026-07-10 收录 2026-07-10 值得跟进

一句话判断

提供了可量化的性能提升数据(+500 Elo,95% 胜率),对评估大模型演示文稿质量的基准测试有价值。

核心信息

GPT-5.6 Sol (max) 在 AA-Briefcase 智能体知识工作基准测试中获得最高 Presentation Elo 分数,较GPT-5.5 (xhigh)提升约500分,预计在对比中胜率约95%。AA-Briefcase综合二元评级检查与分析质量及演示文稿成对比较,测试模型在复杂项目中的真实知识工作任务表现。

原始内容

GPT-5.6 Sol (max) has the highest Presentation Elo of any model in AA-Briefcase AA-Briefcase is our new agentic knowledge work benchmark for testing models on realistic knowledge work tasks in complex projects built by industry experts. Grading in AA-Briefcase combines binary Rubric Checks with Analytical Quality and Presentation pairwise comparison. GPT-5.6 Sol (max) has the highest recorded Presentation Elo - its outputs across various file types, including PowerPoint and Excel, are the most professionally presented of any model. This is a significant jump from the model’s predecessor. GPT-5.6 Sol (max) gains ~500 Presentation Elo points on GPT-5.5 (xhigh) - resulting in a projected win rate of ~95% in head-to-head visual comparisons between the models. ![photo](https://pbs.twimg.com/media/HM4lsdvbIAAInE0.jpg) Artificial Analysis (@ArtificialAnlys): Compare models on AA-Briefcase at https://t.co/RgkI2BmI6R

相关动态

01

Kimi K3 在 SpreadsheetBench 2 上排名第一

Kimi K3 在 SpreadsheetBench 2 基准测试中排名第一,超越 Claude Fable 5,完成 34.8% 的 workflow 任务。该基准测试涉及 321 个专家策划任务,平均每个任务包含 11.8 个工作表和 593.5 个单元格更改,聚焦于完整的电子表格工作簿执行。

twitter关注列表2026-07-18#模型#技术突破#评测
观察
02

Mind Lab 开源长文本强化学习项目

Mind Lab 开源了一个长文本强化学习(RL)项目,支持 2M tokens 的长上下文。相比以往需要数千个 GPU 的 1M 上下文研究,该项目仅需 8 个 GPU 即可完成,大幅降低了超长上下文研究的算力门槛。

twitter关注列表2026-07-17#开源#技术突破#模型
观察
03

Kimi K3非对称竞争分析

Kimi K3在Arena前端/WebDev人类盲测中以1679 Elo排名第一,领先Fable 5(1631)和GPT-5.6 Sol(1618),但在综合智能指数中仅排第四(57.1分),落后Fable 5(59.9)和GPT-5.6 Sol(58.9)。K3定价15美元/百万输出tokens,远低于Fable 5的50美元,并计划7...

twitter关注列表2026-07-18#AI#模型#技术突破
观察
04

二年级学生Jo Nagai发现蝴蝶可遗传记忆

日本东京都二年级学生Jo Nagai注意到养育的食蚜 caterpillars 在蝴蝶化后仍保持对薰衣草的回避行为,经Georgetown大学Entomologist Dr. Martha Weiss合作完成实验:70%训练过的蝴蝶及其后代均表现出对薰衣草的遗传性回避,记忆在全变态发育中保存并遗传。

twitter关注列表2026-07-18#研究#技术突破#信息
值得跟进
05

AI 代码生成大势所趋

Greg Isenberg 发文指出,相比一年前的手写代码实践,如今大多数工程代码已由 AI 生成,标志着编程范式的根本转变。他援引了 Google 75% 新代码由 AI 生成、Anthropic 90%+ 代码由 Claude 编写、GitClear 代码重复率上升 81% 复用率下降 70% 等具体数据,并引用 Dario Amod...

twitter关注列表2026-07-18#技术突破#行业动态#分析
观察