发生了什么
Moonshot AI 发布 Kimi K3(2.8T 参数, 1M 上下文, 16 专家, $3/$15 API)
为什么值得关注
它 introduces a 2.8T sparse expert model with Delta Attention and autonomous chip‑design, delivering up to 6.3× faster decoding and self‑optimizing training—an unprecedented efficiency leap.
信息来源
以下内容来自公开来源,可打开原文继续核验。
Moonshot AI 发布 Kimi K3
Moonshot AI's Kimi K3 is a 2.8‑trillion‑parameter multimodal model with a native 1‑million‑token context window, using a sparse expert system that activates only 16 of its 896 experts per token (about 1.8 % of the pool) and Delta Attention for up to 6.3× faster decoding in million‑token contexts. The model is API‑available at $3/$15 per million input/output tokens and requires roughly 1.4 TB of uncompressed weights, needing about 14‑16 GB10 nodes (≈115 GB per node) or supernodes with 64+ accelerators. In its technical report, Kimi K3 demonstrated autonomous chip design (a 4 mm², 1.46 M‑cell ASIC achieving >8,700 tokens/s), built a GPU compiler (MiniTriton) and self‑optimized training kernels, cutting forward‑plus‑backward time from 283.6 ms to 114.4 ms, completed deep research tasks such as BrowseComp, and reproduced a computational‑astrophysics workflow in ~2 h versus weeks for a researcher, finding inconsistencies in published formulas.
打开原始来源 ↗