Signal Brief

腾讯Hunyuan Hy3成本降低35倍达到Gemini 3.5级物理质量

腾讯新模型Hunyuan Hy3在物理模拟测试中达到Gemini 3.5水平,成本降低35倍。测试显示Hy3仅用29,757 tokens($0.006),远低于Gemini 3.5的23,300 tokens($0.21),而DeepSeek-V4花费50,600 tokens($0.009)但效...

twitter关注列表 Rohan Paul (@rohanpaul_ai) 发布 2026-07-06 收录 2026-07-06 观察

一句话判断

差异点在于具体成本对比数据,以及DeepSeek-V4的意外表现,值得查看原始测试细节。

核心信息

腾讯新模型Hunyuan Hy3在物理模拟测试中达到Gemini 3.5水平,成本降低35倍。测试显示Hy3仅用29,757 tokens($0.006),远低于Gemini 3.5的23,300 tokens($0.21),而DeepSeek-V4花费50,600 tokens($0.009)但效果最差。

原始内容

More good news for local LLMs. Tencent’s new Hunyuan Hy3 reaches Gemini 3.5-level physics quality for 35x less cost. Test was done on atomic[.]chat, a desktop app that runs LLMs locally. The prompt asked 4 models to build bowling, air hockey, and pool simulations. The harder part was preserving physical cause and effect. A strike needs collision timing, mass transfer, pin rotation, friction, and believable scattering. A pool break exposes the same weakness, because every wrong angle compounds immediately. Interestingly, DeepSeek-V4 spent the highest number of tokens (50,600 ), yet produced the weakest visual physics in this test. https://video.twimg.com/amplify_video/2074222891198185472/vid/avc1/1080x1080/wqCUabawMfoY9Tsl.mp4?tag=28 > **引用原帖 atomic.chat (@atomic_chat_hq):** > New Hunyuan Hy3 hits Gemini 3.5 quality on physics for 35x cheaper! > We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos > Prompts: > - A bowling ball knocking down the pins > - An air hockey rally that ends in a goal > - A pool break scattering the rack > Outputs: > Hunyuan Hy3: 29,757 tokens, $0.006 > Gemini 3.5: 23,300 tokens, $0.21 > GLM-5.2: 25,454 tokens, $0.07 > DeepSeek-V4: 50,600 tokens, $0.009 > Tencent's Hy3 matched Gemini across all three: clean collisions, the puck bounced true, the pins scattered like a real strike, the rack broke with real momentum, nothing clipped or floated. GLM is genuinely strong on pure coding tasks, but the moment the job steps outside clean code it gives way. DeepSeek was the letdown, it burned the most tokens of anyone (50k, almost 2x Hy3) and still turned in the weakest scenes > https://x.com/atomic_chat_hq/status/2074202885517443364 Rohan Paul (@rohanpaul_ai): Github of Atomic-Chat. "an open source alternative to ChatGPT that runs 100% offline on your computer." https://t.co/rfprzJzHRO Download it here. https://t.co/aMAZoaXjZ4

相关动态

01

Kimi K3 在 SpreadsheetBench 2 上排名第一

Kimi K3 在 SpreadsheetBench 2 基准测试中排名第一,超越 Claude Fable 5,完成 34.8% 的 workflow 任务。该基准测试涉及 321 个专家策划任务,平均每个任务包含 11.8 个工作表和 593.5 个单元格更改,聚焦于完整的电子表格工作簿执行。

twitter关注列表2026-07-18#模型#技术突破#评测
观察
02

Mind Lab 开源长文本强化学习项目

Mind Lab 开源了一个长文本强化学习(RL)项目,支持 2M tokens 的长上下文。相比以往需要数千个 GPU 的 1M 上下文研究,该项目仅需 8 个 GPU 即可完成,大幅降低了超长上下文研究的算力门槛。

twitter关注列表2026-07-17#开源#技术突破#模型
观察
03

Kimi K3非对称竞争分析

Kimi K3在Arena前端/WebDev人类盲测中以1679 Elo排名第一,领先Fable 5(1631)和GPT-5.6 Sol(1618),但在综合智能指数中仅排第四(57.1分),落后Fable 5(59.9)和GPT-5.6 Sol(58.9)。K3定价15美元/百万输出tokens,远低于Fable 5的50美元,并计划7...

twitter关注列表2026-07-18#AI#模型#技术突破
观察
04

二年级学生Jo Nagai发现蝴蝶可遗传记忆

日本东京都二年级学生Jo Nagai注意到养育的食蚜 caterpillars 在蝴蝶化后仍保持对薰衣草的回避行为,经Georgetown大学Entomologist Dr. Martha Weiss合作完成实验:70%训练过的蝴蝶及其后代均表现出对薰衣草的遗传性回避,记忆在全变态发育中保存并遗传。

twitter关注列表2026-07-18#研究#技术突破#信息
值得跟进
05

AI 代码生成大势所趋

Greg Isenberg 发文指出,相比一年前的手写代码实践,如今大多数工程代码已由 AI 生成,标志着编程范式的根本转变。他援引了 Google 75% 新代码由 AI 生成、Anthropic 90%+ 代码由 Claude 编写、GitClear 代码重复率上升 81% 复用率下降 70% 等具体数据,并引用 Dario Amod...

twitter关注列表2026-07-18#技术突破#行业动态#分析
观察