Signal Brief

PACE: A Proxy for Agentic Capability Evaluation

Yueqi Song 及合作者发布 PACE,一种通过从非代理基准中选取少量实例来预测代理基准性能的方法。在 14 个模型、4 个代理基准和 19 个非代理基准上,PACE 实现了 3.80% 的绝对分数预测 MAE、0.81 的 Spearman 排名相关、约 84% 的成对偏好准确率,以及约 1...

twitter关注列表 马东锡 NLP (@dongxi_nlp) 发布 2026-07-06 收录 2026-07-06 观察

一句话判断

该方法揭示了代理基准所需的不同能力(如规划、验证),并提供了低成本的评估替代方案,值得阅读原文了解具体实现。

核心信息

Yueqi Song 及合作者发布 PACE,一种通过从非代理基准中选取少量实例来预测代理基准性能的方法。在 14 个模型、4 个代理基准和 19 个非代理基准上,PACE 实现了 3.80% 的绝对分数预测 MAE、0.81 的 Spearman 排名相关、约 84% 的成对偏好准确率,以及约 100 倍的成本降低。

原始内容

PACE: A Proxy for Agentic Capability Evaluation > **引用原帖 Yueqi Song (@yueqi_song):** > 🚀Excited to release PACE: A Proxy for Agentic Capability Evaluation! > Evaluating LLM agents on benchmarks like SWE-Bench and GAIA is expensive, slow, and infrastructure-heavy, often costing $$$ and taking hours or days per model. > ❓But do we always need to run full agentic evaluations? > In PACE, we show that agentic benchmark performance can be accurately predicted from a small, carefully selected set of cheap non-agentic benchmark instances. > PACE automatically selects proxy instances from existing benchmarks covering skills like instruction following, planning, tool use, reasoning, coding, retrieval, and multimodal understanding. > Across 14 models, 4 agentic benchmarks, and 19 non-agentic benchmarks, PACE-BENCH achieves: > ✅ 3.80% MAE for absolute score prediction > ✅ 0.81 Spearman correlation for model ranking > ✅ ~84% pairwise preference accuracy > ✅ ~100× lower cost than target benchmark sampling > Beyond prediction, PACE also reveals what capabilities different agentic benchmarks actually require, e.g., planning, verification, long-context aggregation, and instruction following. > We hope PACE makes agentic evaluation cheaper, faster, and more accessible for model development, model selection, and routing :) > 📃 Paper: https://t.co/I6iYUCADov > 💻 Code: https://t.co/tLd7HrSgEA > I'm incredibly grateful to have worked with @lintangsutawika, @Jiarui_Liu_, @lltjuatja, @JiayiiGeng, @lrzneedresearch, @daniel_js_lee, @Aditya_Soni_8, @Vincent92965015, @xiangyue96, and @gneubig . > https://x.com/yueqi_song/status/2074180763302670648

相关动态

02

Kimi K3非对称竞争分析

Kimi K3在Arena前端/WebDev人类盲测中以1679 Elo排名第一,领先Fable 5(1631)和GPT-5.6 Sol(1618),但在综合智能指数中仅排第四(57.1分),落后Fable 5(59.9)和GPT-5.6 Sol(58.9)。K3定价15美元/百万输出tokens,远低于Fable 5的50美元,并计划7...

twitter关注列表2026-07-18#AI#模型#技术突破
观察
03

二年级学生Jo Nagai发现蝴蝶可遗传记忆

日本东京都二年级学生Jo Nagai注意到养育的食蚜 caterpillars 在蝴蝶化后仍保持对薰衣草的回避行为,经Georgetown大学Entomologist Dr. Martha Weiss合作完成实验:70%训练过的蝴蝶及其后代均表现出对薰衣草的遗传性回避,记忆在全变态发育中保存并遗传。

twitter关注列表2026-07-18#研究#技术突破#信息
值得跟进
05

BestBlogs 早报 · 07-18

月之暗面发布Kimi K3,2.8万亿参数,896选16的Stable LatentMoE,上下文100万token,接近Fable-5但仍落后最强闭源模型,完整权重7月27日前开源;VentureBeat调查显示54%企业已发生AI代理安全事件;xAI开源Grok Build(84万行Rust代码)并残留上传用户代码痕迹;Cursor评...

twitter关注列表2026-07-17#AI#模型发布#开源
观察