Signal Brief

Tencent Hy3 Agent Arena排名#25

Arena.ai评测显示,腾讯Hy3模型在Agent Arena中排名第25位(开放权重模型第5位),在Frontend Code Arena中排名第16位(开放模型第2位)。Hy3在工具使用(bash错误恢复,+2.6%)方面表现强劲,但在可操纵性(用户推回时调整能力)方面较弱,得分为-7.1%。

twitter关注列表 Tencent Hy (@TencentHunyuan) 发布 2026-07-22 收录 2026-07-22 观察

一句话判断

提供了Hy3在真实Agent任务中的详细优劣数据,对评估开源Agent模型性能有参考价值。

核心信息

Arena.ai评测显示,腾讯Hy3模型在Agent Arena中排名第25位(开放权重模型第5位),在Frontend Code Arena中排名第16位(开放模型第2位)。Hy3在工具使用(bash错误恢复,+2.6%)方面表现强劲,但在可操纵性(用户推回时调整能力)方面较弱,得分为-7.1%。

原始内容

Agent & Coding 🔥🔥🔥 > **引用原帖 Arena.ai (@arena):** > Hy3 by Tencent is #5 in Agent Arena for open-weight models (#25 overall)! It also ranks as the #2 open model in the Frontend Code Arena (#16 overall)! > In Agent Arena: Hy3 lands at #25 overall (net -2.2%). Hy3 has strengths in tool-use (recovering well from CLI/bash errors, +2.6% and #25), its biggest weakness is steerability as it struggles to course-correct when users push back, coming in at -7.1% (#30). > Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents. > We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model. > Below we break down how Hy3 scored across 5 key signals, drawn from tasks submitted by a global community of users. Here’s an overview on the signals: > User-satisfaction proxies > - Confirmed Success: an explicit "yes that worked" feedback from the user > - Praise vs. Complaint: implicit sentiment in users reactions > - Steerability: can the model course-correct when you push back? > Tool-use proxies > - Bash Recovery: how it recovers from CLI errors (primary signal for tool use) > - Tool Hallucination: does it call tools that don't exist > Congrats to the @TencentHunyuan team on this release! > https://x.com/arena/status/2079698021085016270

相关动态

01

AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?

研究人员提出了 AutoLab 基准测试,包含 36 项涉及系统加速、谜题、模型开发及 CUDA kernel 工作的任务,旨在评估模型在长时程自动研究中的表现。通过对 17 个主流模型的测试发现,成功的关键不在于初始想法的质量,而在于模型的持续测试与反馈迭代能力;其中 Claude Opus 4.6 在该基准测试中表现最佳,原因在于其能...

twitter关注列表2026-07-22#技术#模型#研究
观察
02

Kimi K3 30分钟构建VR伴侣

Kimi团队展示Kimi K3模型在收到一个3D模型和一条提示后,于30分钟内自主构建了一个VR伴侣,该伴侣能够听、回应、改变表情并在场景间移动,期间K3独立发现并部署了本地ASR和TTS连接LLM,并构建了VR场景。

twitter关注列表2026-07-22#AI#模型#多模态
观察
03

dots-note-3.0获IMO满分金牌

小红书dots团队的大模型dots-note-3.0在第67届IMO中获得42分满分,超过金牌线13分,成为全球首个在IMO官方评卷中斩获满分的大模型,也是中国首个IMO金牌大模型。此前Gemini Deep Think曾获35分。该模型采用智能体化推理系统和加强归纳法,并计划近期开源。

twitter关注列表2026-07-22#模型#技术突破#AI模型
值得跟进
04

Anthropic 用 Claude Code 实现百万行代码迁移新模式

Anthropic 团队使用 Claude Code 将 Zig 项目迁移到 Rust,百万行代码不到两周完成,100% 测试通过,仅 19 个回归问题全部修复,成本约 16.5 万美元。文章介绍了六步迁移流程,强调以强 judge 和对抗 agent 验证测试断言,可复制的工厂流水线模板将语言级重构门槛降至小团队可行。

twitter关注列表2026-07-22#技术#模型发布#开源
值得跟进