Signal Brief

Arena 增加事实评分

Arena 在其排行榜中加入了基于真实用户对话的事实核查评分,对模型的可靠性进行衡量。他们从超过 200 万条模型在真实对话中提出的可验证主张中抽样验证,覆盖约 13 万文本竞赛和 4 万搜索竞赛。启用事实权重后,GPT‑5.5 在文本竞赛中上升 13 名,Muse Spark 下降 13 名,Cl...

twitter关注列表 meng shao (@shao__meng) 发布 2026-07-15 收录 2026-07-16 观察

一句话判断

事实评分显著重新排列了模型排名,尤其是 GPT‑5.5 上升 13 名

核心信息

Arena 在其排行榜中加入了基于真实用户对话的事实核查评分,对模型的可靠性进行衡量。他们从超过 200 万条模型在真实对话中提出的可验证主张中抽样验证,覆盖约 13 万文本竞赛和 4 万搜索竞赛。启用事实权重后,GPT‑5.5 在文本竞赛中上升 13 名,Muse Spark 下降 13 名,Claude Fable 5 略降,而在搜索竞赛中 GPT‑5.5-search 登顶,Gemini grounding 跌至第 13 名。

原始内容

meng shao (@shao__meng) 转发了 Sumanth (@Sumanth_077) 的帖子: The model humans prefer the most is not always the most accurate one! Arena just added factuality scoring to their leaderboard alongside human preference. The idea: human preference measures whether a response felt good to read. Factuality measures the correctness of claims in a model's response. These are two different things. Most benchmarks test models on fixed, predetermined questions in a lab setting. What Arena is doing differently: measuring factuality in real, open-ended user conversations at scale. Over 2 million claims extracted from actual conversations, verified against the web, across 130k Text Arena battles and 40k Search Arena battles. The way it works: they randomly sample battles, extract web-verifiable claims from each model response, verify them against the web, and calculate a factuality score per model. The final ranking is a weighted combination of human preference and factuality score. When factuality weight is enabled, the shifts are significant. In Text Arena, GPT-5.5 jumps 13 spots. Muse Spark drops 13. Claude Fable 5 moves down slightly. In Search Arena, GPT-5.5-search takes the top spot while Gemini grounding drops from 7 to 13. I've shared the full methodology blog in the replies! ![photo](https://pbs.twimg.com/media/HNTHC3paIAAksR2.jpg) > **引用原帖 Arena.ai (@arena):** > Introducing factuality in the Arena: a new ranking of models according to a weighted combination of human preference and factuality. > Model rankings are now viewable according to a weighted combination of human preference and factuality. Factuality is live in our Text and Search Arenas as a non-default toggle. > We audit model responses by randomly sampling battles and extracting web-verifiable claims. We then verify these claims and compare the average correctness between model responses. > To power these rankings, we’ve labeled over 2 million claims made by LLMs in real-world conversations, 1.3+ million from Text Arena, and 700k+ from Search Arena. > Notable highlights with factuality enabled in the Text Arena: > - Claude Fable 5 moves down slightly to spot #2 > - GPT-5.5 saw the largest increase, moving up 13 spots into the #7 spot > - Muse Spark dropped the most from #7 to #20 (-13pt) > By labs, Meta saw the largest drop from #2 to #5, while Anthropic overall held the #1 spot. Looking at only open model providers, Xiaomi saw the largest improvement, jumping from #9 to #6. > Learn more about the findings and methodology in this thread. > https://x.com/arena/status/2077432293023678685

相关动态

01

Kimi K3 在 SpreadsheetBench 2 上排名第一

Kimi K3 在 SpreadsheetBench 2 基准测试中排名第一,超越 Claude Fable 5,完成 34.8% 的 workflow 任务。该基准测试涉及 321 个专家策划任务,平均每个任务包含 11.8 个工作表和 593.5 个单元格更改,聚焦于完整的电子表格工作簿执行。

twitter关注列表2026-07-18#模型#技术突破#评测
观察
02

Mind Lab 开源长文本强化学习项目

Mind Lab 开源了一个长文本强化学习(RL)项目,支持 2M tokens 的长上下文。相比以往需要数千个 GPU 的 1M 上下文研究,该项目仅需 8 个 GPU 即可完成,大幅降低了超长上下文研究的算力门槛。

twitter关注列表2026-07-17#开源#技术突破#模型
观察
03

Kimi K3非对称竞争分析

Kimi K3在Arena前端/WebDev人类盲测中以1679 Elo排名第一,领先Fable 5(1631)和GPT-5.6 Sol(1618),但在综合智能指数中仅排第四(57.1分),落后Fable 5(59.9)和GPT-5.6 Sol(58.9)。K3定价15美元/百万输出tokens,远低于Fable 5的50美元,并计划7...

twitter关注列表2026-07-18#AI#模型#技术突破
观察
04

二年级学生Jo Nagai发现蝴蝶可遗传记忆

日本东京都二年级学生Jo Nagai注意到养育的食蚜 caterpillars 在蝴蝶化后仍保持对薰衣草的回避行为,经Georgetown大学Entomologist Dr. Martha Weiss合作完成实验:70%训练过的蝴蝶及其后代均表现出对薰衣草的遗传性回避,记忆在全变态发育中保存并遗传。

twitter关注列表2026-07-18#研究#技术突破#信息
值得跟进
05

AI 代码生成大势所趋

Greg Isenberg 发文指出,相比一年前的手写代码实践,如今大多数工程代码已由 AI 生成,标志着编程范式的根本转变。他援引了 Google 75% 新代码由 AI 生成、Anthropic 90%+ 代码由 Claude 编写、GitClear 代码重复率上升 81% 复用率下降 70% 等具体数据,并引用 Dario Amod...

twitter关注列表2026-07-18#技术突破#行业动态#分析
观察