Signal Brief

Artificial Analysis 评测 Thinking Machines Lab 的 Inkling 模型

Artificial Analysis 评测 Thinking Machines Lab 的 Inkling 模型在 AA-Briefcase 基准上获得 Elo 836,Rubric 得分 19.3%,低于 MiMo-V2.5-Pro(21.4%)但高于 DeepSeek V4 Flash max...

twitter关注列表 Artificial Analysis (@ArtificialAnlys) 发布 2026-07-22 收录 2026-07-22 观察

一句话判断

新增了 Inkling 在复杂知识工作基准上的具体表现,包括多轮少工具调用的独特策略,值得关注其效率特点。

核心信息

Artificial Analysis 评测 Thinking Machines Lab 的 Inkling 模型在 AA-Briefcase 基准上获得 Elo 836,Rubric 得分 19.3%,低于 MiMo-V2.5-Pro(21.4%)但高于 DeepSeek V4 Flash max(18.7%)和 Gemini 3.5 Flash-Lite(14.8%);Presentation Elo 863,Analytical Quality Elo 764;平均输出 52K token/任务,平均 81 轮/任务但工具调用较少。

原始内容

Thinking Machines Lab’s Inkling scores an Elo of 836 on on our agentic knowledge work benchmark AA-Briefcase, ahead of DeepSeek V4 Flash but below leading open weights models including Nemotron 3 Ultra and GLM-5.2 Our new agentic knowledge work benchmark, AA-Briefcase, tests models on realistic tasks across thousands of input files, requiring deliverables such as spreadsheets, presentations, and UI mock-ups. Model performance is measured across three dimensions: binary rubric checks for ground-truth correctness, pairwise grading on analytical quality, and pairwise grading on presentation quality. The AA-Briefcase Elo is a single metric that combines results across all three dimensions Key Takeaways: ➤ Scores 19.3% on the AA-Briefcase rubric, below MiMo-V2.5-Pro (21.4%) but above DeepSeek V4 Flash max (18.7%) and Gemini 3.5 Flash-Lite (14.8%). Inkling’s rubric score is worst on tasks which include non standard “Other” file types (i.e., not Excel, PowerPoint, PDF, or Word), despite its native multimodal support ➤ Performs higher on Presentation than Analytical Quality with Elo scores of 863 and 764 respectively. Presentation and Analytical Quality are both measured using separate, independent pairwise checks from model submissions. Graders compare two submissions for the same task and pick the one that is more professionally presented (Presentation) and the one with deeper, better-structured analysis (Analytical Quality) ➤ Uses 52K output tokens on average per AA-Briefcase task and 5M output tokens for the full suite, slightly more than models with similar scores on AA-Briefcase. Inkling uses the most tokens for Excel deliverables, followed by Word, PDF, PowerPoint, and Other types ➤ Has one of the highest mean turns per AA-Briefcase task (81) with also one of the widest ranges, resulting in a much lower median (49). Despite one of the higher average turns per task, Inkling comparatively uses fewer tool calls per turn on average (0.5) ![photo](https://pbs.twimg.com/media/HN3E_53XoAAUoxH.jpg) Artificial Analysis (@ArtificialAnlys): Despite one of the higher average turns per task, Inkling comparatively uses fewer tool calls per turn on average (0.5) https://t.co/se8nCsu1JU Artificial Analysis (@ArtificialAnlys): For more details: https://t.co/RgkI2Bmahj

相关动态

02

梁文锋投资人会议语录

DeepSeek创始人梁文锋在投资人会议上分享对AGI路线、开源、竞争等观点,认为AGI的关键阶梯是CoT→Agent→持续学习,多模态和视频生成不是智能本质,开源和低利润是战略,国内模型公司会收敛为2-3家。

twitter关注列表2026-07-22#AI#技术#行业动态
观察
03

利用 Embedding 模型实现图像打标

Han Xiao 通过将冻结的 jina-v5-omni 多模态嵌入模型进行测试时缩放(scaled at test time),在无需训练、无需第二模型及外部知识的硬约束下,实现了强大的开放词汇多标签 n-gram 图像打标器。

twitter关注列表2026-07-22#技术#研究#模型
观察
04

AI 30年图论猜想被推翻

Dmitry Rybin 通过 GPT 5.6 Pro 发现 Dinitz-Garg-Goemans 猜想错误,该猜想开放约30年。具体反例图显示分式流成本58,不可分裂流(容量违反≤15)成本至少60。

twitter关注列表2026-07-22#AI#技术突破#研究
观察