Signal Brief

vLLM 发布 AFD 插件实现 MoE 推理优化

vLLM 发布实验性插件 vLLM AFD Plugin,实现 Attention-FFN Disaggregation (AFD),将 MoE 模型的注意力路径与专家路径分离为独立服务,支持独立扩缩,兼容 NVIDIA GPU 和 Ascend NPU。

twitter关注列表 Hao AI Lab (@haoailab) 发布 2026-07-24 收录 2026-07-24 观察

一句话判断

首次将 AFD 技术集成到 vLLM,可独立扩缩 attention 和 expert,对 MoE 推理效率有实际优化。

核心信息

vLLM 发布实验性插件 vLLM AFD Plugin,实现 Attention-FFN Disaggregation (AFD),将 MoE 模型的注意力路径与专家路径分离为独立服务,支持独立扩缩,兼容 NVIDIA GPU 和 Ascend NPU。

原始内容

Hao AI Lab (@haoailab) 转发了 Roger Wang (@rogerw0108) 的帖子: The idea of AFD has been out for a while and the teams were finally able to put it in practice! We're also grateful for the valuable feedbacks from @FuYichao123, @Yuxuan_Zhang13, @Junda_Chen_ and others from @haoailab (Check out FastAFD!) and welcome more collaborations! > **引用原帖 vLLM (@vllm_project):** > 🎉 Congrats to the teams behind vLLM AFD Plugin (Ascend & vLLM, @StepFun_ai, @AntGroup, FastAFD): a new experimental plugin under vllm-project that brings Attention-FFN Disaggregation to MoE serving. > Attention and the expert/FFN path are two very different workloads that normally share one topology. AFD runs them as separate services, so you can scale attention and experts independently. Same vLLM serving surface, no fork. @NVIDIA GPU and Ascend NPU. > https://x.com/vllm_project/status/2080673192939807083

相关动态

02

OpenCode爆炸增长数据

Y Combinator 在播客中披露,OpenCode(开源替代 Claude Code 和 Codex 的 AI 编码工具)自年初起增长至 460 万周活用户、1300 万月活用户,年化收入约 4000 万美元,处理了 7 万亿 tokens,CEO Jay V 谈到 Anthropic 争议和 20 倍增长驱动因素。

twitter关注列表2026-07-24#行业动态#AI#开源
观察
03

Atomic Agent GAIA基准胜Hermes

Atomic Agent在GAIA Level 1基准测试中以69.8%正确率超越Hermes的58.5%,速度快1.6倍(3h12m vs 5h10m),使用相同4-bit Qwen-3.6-35B模型和Apple M4 Max硬件,开源MIT许可。

twitter关注列表2026-07-24#开源#评测#技术突破
观察
04

Grok 集成 Google Workspace

Grok 已集成到 Google Workspace,通过侧边栏插件在 Sheets、Slides、Docs 中提供数据解释、公式书写、演示生成、文档起草等功能,无需频繁切换标签页,组织可批量部署。

twitter关注列表2026-07-24#产品发布#大模型#AI
观察