Signal Brief

Moonshot AI 发布 Kimi K3

Moonshot AI's Kimi K3 is a 2.8‑trillion‑parameter multimodal model with a native 1‑million‑token context window, using a sparse expert system that act...

twitter关注列表 Rohan Paul (@rohanpaul_ai) 发布 2026-07-16 收录 2026-07-17 值得跟进

一句话判断

它 introduces a 2.8T sparse expert model with Delta Attention and autonomous chip‑design, delivering up to 6.3× faster decoding and self‑optimizing training—an unprecedented efficiency leap.

核心信息

Moonshot AI's Kimi K3 is a 2.8‑trillion‑parameter multimodal model with a native 1‑million‑token context window, using a sparse expert system that activates only 16 of its 896 experts per token (about 1.8 % of the pool) and Delta Attention for up to 6.3× faster decoding in million‑token contexts. The model is API‑available at $3/$15 per million input/output tokens and requires roughly 1.4 TB of uncompressed weights, needing about 14‑16 GB10 nodes (≈115 GB per node) or supernodes with 64+ accelerators. In its technical report, Kimi K3 demonstrated autonomous chip design (a 4 mm², 1.46 M‑cell ASIC achieving >8,700 tokens/s), built a GPU compiler (MiniTriton) and self‑optimized training kernels, cutting forward‑plus‑backward time from 283.6 ms to 114.4 ms, completed deep research tasks such as BrowseComp, and reproduced a computational‑astrophysics workflow in ~2 h versus weeks for a researcher, finding inconsistencies in published formulas.

原始内容

Beijing-based Moonshot AI's Kimi K3 just dropped. - 1 Mn token context window, natively multimodal. (i.e. 750,000 words of code or documentation in a single prompt.) - Total 2.8 trillion params, but activates only 16 of its 896 experts at a time. That is roughly 1.8% of the expert pool per token. - Some of the benchmarks put it Opus 4.8/ GPT 5.6 Sol / Fable 5 territory. - Its Delta Attention enables up to 6.3x faster decoding in million-token contexts - available via API at $3/$15 per million input/output tokens - K3’s 2.8T MXFP4 weights require roughly 1.4 TB before quantization metadata and runtime overhead. At an assumed 115 GB usable per GB10 node (NVIDIA’s compact Grace Blackwell AI chip), 14–16 nodes is expected to hold the raw weights realistically. Including activations, KV cache and runtime memory etc. Moonshot itself recommends supernodes containing 64 or more accelerators. Some very cool findings from technical report - K3 autonomously designed, optimized, and verified a working AI chip in a single 48-hour run—specifically to serve a smaller model built on K3’s own architecture. The simulated chip reportedly reached 8,700+ tokens/second, contained 1.46 million standard cells, and fit within 4 mm². - An early version of K3 handled the majority of the kernel-optimization work used to develop K3 itself. - K3 built a GPU compiler from scratch. created MiniTriton, optimization passes, PTX generation, and runtime. Matched or beat Triton on some workloads and successfully trained nanoGPT end to end. - In a 15-hour autonomous run, Kimi K3 redesigned a production-scale training kernel and cut forward-plus-backward time from 283.6 ms to 114.4 ms. - For a 42-year semiconductor-industry report, Kimi says K3 performed 2,800+ web searches/fetches, 1,100+ terminal data pulls, processed 11,000+ pages, and recursively improved the work over 120+ rounds. - K3 reproduced a computational-astrophysics research workflow in roughly two hours, versus an estimated one to two weeks for an experienced researcher. It reviewed 20+ papers, evaluated 300+ equations of state, found inconsistencies in published formulas, and wrote 3,000+ lines of Python. ![photo](https://pbs.twimg.com/media/HNYqg7daEAA0hdm.png) > **引用原帖 Kimi.ai (@Kimi_Moonshot):** > Introducing Kimi K3: Open Frontier Intelligence > 🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal > 🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts > 🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost > 🔹 Built for long-horizon agentic coding and self-evolving workflows > Kimi K3 is now live on on https://t.co/zrk6zZxZUo, Kimi Work, Kimi Code, and the Kimi API. > Open Weights by July 27, 2026. > 🔗 API: https://t.co/XCrgjXAqMw > 🔗 Tech blog: https://t.co/YTfiMSNM1f > https://x.com/Kimi_Moonshot/status/2077830229968683203 Rohan Paul (@rohanpaul_ai): Kimi K3 appears to beat every model shown on BrowseComp while costing only a fraction as much per task. Its claiming a dramatically better intelligence-per-dollar tradeoff. BrowseComp assesses the capability of an AI agent to repeatedly search the internet and discover obscure, hard-to-find facts that require multiple searches and validation. i.e. a high BrowseComp score means you are good at deep research and web investigation, but not necessarily better general intelligence, writing, coding, or everyday browsing.

相关动态

01

Kimi K3 在 SpreadsheetBench 2 上排名第一

Kimi K3 在 SpreadsheetBench 2 基准测试中排名第一,超越 Claude Fable 5,完成 34.8% 的 workflow 任务。该基准测试涉及 321 个专家策划任务,平均每个任务包含 11.8 个工作表和 593.5 个单元格更改,聚焦于完整的电子表格工作簿执行。

twitter关注列表2026-07-18#模型#技术突破#评测
观察
02

Mind Lab 开源长文本强化学习项目

Mind Lab 开源了一个长文本强化学习(RL)项目,支持 2M tokens 的长上下文。相比以往需要数千个 GPU 的 1M 上下文研究,该项目仅需 8 个 GPU 即可完成,大幅降低了超长上下文研究的算力门槛。

twitter关注列表2026-07-17#开源#技术突破#模型
观察
03

Kimi K3非对称竞争分析

Kimi K3在Arena前端/WebDev人类盲测中以1679 Elo排名第一,领先Fable 5(1631)和GPT-5.6 Sol(1618),但在综合智能指数中仅排第四(57.1分),落后Fable 5(59.9)和GPT-5.6 Sol(58.9)。K3定价15美元/百万输出tokens,远低于Fable 5的50美元,并计划7...

twitter关注列表2026-07-18#AI#模型#技术突破
观察
04

二年级学生Jo Nagai发现蝴蝶可遗传记忆

日本东京都二年级学生Jo Nagai注意到养育的食蚜 caterpillars 在蝴蝶化后仍保持对薰衣草的回避行为,经Georgetown大学Entomologist Dr. Martha Weiss合作完成实验:70%训练过的蝴蝶及其后代均表现出对薰衣草的遗传性回避,记忆在全变态发育中保存并遗传。

twitter关注列表2026-07-18#研究#技术突破#信息
值得跟进
05

AI 代码生成大势所趋

Greg Isenberg 发文指出,相比一年前的手写代码实践,如今大多数工程代码已由 AI 生成,标志着编程范式的根本转变。他援引了 Google 75% 新代码由 AI 生成、Anthropic 90%+ 代码由 Claude 编写、GitClear 代码重复率上升 81% 复用率下降 70% 等具体数据,并引用 Dario Amod...

twitter关注列表2026-07-18#技术突破#行业动态#分析
观察