Signal Brief

Kimi K3 以更低成本胜 Err Opus 4.8 生成复古游戏

Atomic Chat 的本地 LLM 桌面应用给 Kimi K3、GPT‑5.6 与 Opus 4​ន三种模型同样任务,要求各自生成 Road Fighter、Battle City、Q*bert 三款自玩型 HTML 版复古游戏,包含图形、敌人фа和完整 AI 玩家。输出结果显示 Kimi K3...

twitter关注列表 Rohan Paul (@rohanpaul_ai) 发布 2026-07-16 收录 2026-07-17 观察

一句话判断

Kimi K3 在相同任务下以更低代价提供更完整游戏逻辑,值得进一步复现验证。

核心信息

Atomic Chat 的本地 LLM 桌面应用给 Kimi K3、GPT‑5.6 与 Opus 4​ន三种模型同样任务,要求各自生成 Road Fighter、Battle City、Q*bert 三款自玩型 HTML 版复古游戏,包含图形、敌人фа和完整 AI 玩家。输出结果显示 Kimi K3 共 18.4K tokens,成本 $0.28;GPT‑5.6 18.1K tokens,$0.28;Opus 4.8 21.3K tokens,$0.54。Kimi K3 的 Q*bert 通过改变立方体状态、绕庙山移动并躲避敌人实现完整游戏循环;Opus 4.8 在大多数功能上存在错误,GPT‑5.6 虽有更好纹理但车/坦克逻辑不完整。

原始内容

Kimi K3 just beat Opus 4.8 at building retro games for half the cost. Atomic Chat, a desktop app that runs LLMs locally, asked each model to build Road Fighter, Battle City, and Q*bert as self-playing HTML files. To make each result work, the graphics, enemies, game rules and a player that could play on their own all had to be in the same file. Because a single error can ruin the entire simulation, setup tests are more important than code generation. Outputs: Kimi K3: 18.4K tokens, $0.28 GPT-5.6: 18.1K tokens, $0.28 Opus 4.8: 21.3K tokens, $0.54 GPT-5.6 generated the best looking textures, but the logic for the cars/tanks did not work. Shows how great looks can be deceiving for poor control and state management. Kimi's best work was Q*bert, where the agent changed the state of the cube, moved around the pyramid, and avoided an enemy, without breaking the game loop. https://video.twimg.com/amplify_video/2077887033461395456/vid/avc1/1600x900/iRXtGLsK9n9sGGZx.mp4?tag=29 > **引用原帖 atomic.chat (@atomic_chat_hq):** > New Kimi K3 beat Opus 4.8 at making retro games for 2x cheaper! > We gave three models the same task: make three old arcade games that play by themselves. Each game is one HTML file with a bot that plays it > Prompts: > – Road Fighter > – Battle City > – Q*bert > Outputs: > Kimi K3: 18.4K tokens, $0.28 > GPT-5.6: 18.1K tokens, $0.28 > Opus 4.8: 21.3K tokens, $0.54 > Kimi K3 made the best games. Its Q*bert is the best of all three. The player jumps on the cubes, paints them and runs from the purple snake. Its Battle City looks just like the real game and plays right. The tanks move, aim, shoot and break the walls. GPT-5.6 could not handle the cars. Its Battle City broke too. The tank died on its own and the base got hit. But GPT had the nicest textures of the three > https://x.com/atomic_chat_hq/status/2077854291058950261 Rohan Paul (@rohanpaul_ai): Github of Atomic-Chat. "an open source alternative to ChatGPT that runs 100% offline on your computer." https://t.co/rfprzJzHRO Download it here. https://t.co/aMAZoaXjZ4

相关动态

01

Kimi K3 在 SpreadsheetBench 2 上排名第一

Kimi K3 在 SpreadsheetBench 2 基准测试中排名第一,超越 Claude Fable 5,完成 34.8% 的 workflow 任务。该基准测试涉及 321 个专家策划任务,平均每个任务包含 11.8 个工作表和 593.5 个单元格更改,聚焦于完整的电子表格工作簿执行。

twitter关注列表2026-07-18#模型#技术突破#评测
观察
02

BestBlogs 早报 · 07-18

月之暗面发布Kimi K3,2.8万亿参数,896选16的Stable LatentMoE,上下文100万token,接近Fable-5但仍落后最强闭源模型,完整权重7月27日前开源;VentureBeat调查显示54%企业已发生AI代理安全事件;xAI开源Grok Build(84万行Rust代码)并残留上传用户代码痕迹;Cursor评...

twitter关注列表2026-07-17#AI#模型发布#开源
观察
03

Kimi K3 编码代理评测

Artificial Analysis发布Kimi K3在编码代理指数上的评测结果:得分57,排名第5,性能与GPT-5.6 Terra和GPT-5.5持平(57),超过Opus 4.8(55),略低于Grok 4.5(58)和Fable 5(59);成本平均3.18美元/任务,比GPT-5.6 Sol便宜55%。

twitter关注列表2026-07-17#AI模型#评测
观察
05

Runway Agent 在第三方评测中全面领先

Physion Labs 发布 Physion-Arc 1.0 基准测试,对 Runway、Luma、MiniMax、Kling、Utopia 和 TapNow 六款 AI 视频代理进行独立人类评估,使用 30 个电影提示和 16 个指标。Runway Agent 2.0 在叙事连贯性、电影语言和制作质量三项核心维度全部排名第一。

twitter关注列表2026-07-17#评测#多模态#模型
观察