← 返回精选动态
精选动态

MiniMax Sparse Attention

[MiniMax] 发表 [Sparse Attention](28.4x compute reduction, 14.2x faster prefill, 7.6x faster decoding)

更新于 2026年06月16日 05:52 1 条公开来源

发生了什么

[MiniMax] 发表 [Sparse Attention](28.4x compute reduction, 14.2x faster prefill, 7.6x faster decoding)

为什么值得关注

与全注意力相比,MSA 通过架构内的可训练选择器动态选取相关键值块,实现数量级加速,值得阅读原文了解路由设计细节。

信息来源

以下内容来自公开来源,可打开原文继续核验。

01

MiniMax Sparse Attention

twitter关注列表 2026年06月16日 04:39

MiniMax 提出 Sparse Attention(MSA),在 H800 GPU 上实现了 attention 计算量减少 28.4 倍、prefill 加速 14.2 倍、decoding 加速 7.6 倍,同时保持与全注意力版本相当的基准性能。该方法在 Grouped Query Attention 基础上增加小型路由分支,让每个查询组选择要检查的 key-value 块,主分支仅在这些选定块内执行精确注意力。

打开原始来源 ↗
← 返回精选动态