Signal Brief

SANA-Video 2.0 发布

Enze Xie 团队发布 SANA-Video 2.0,采用混合线性-softmax 注意力、块注意力残差和 Sol-Engine 加速,统一 5B 和 14B 模型,在 16 节点 H100 上训练 5B 模型,VBench 总分 84.30,单 H100 生成 720p/5s 仅需 13.06...

twitter关注列表 Emad (@EMostaque) 发布 2026-07-24 收录 2026-07-24 观察

一句话判断

该模型在极低训练资源下实现了高性能视频生成,混合注意力架构值得关注。

核心信息

Enze Xie 团队发布 SANA-Video 2.0,采用混合线性-softmax 注意力、块注意力残差和 Sol-Engine 加速,统一 5B 和 14B 模型,在 16 节点 H100 上训练 5B 模型,VBench 总分 84.30,单 H100 生成 720p/5s 仅需 13.06s,速度比 Wan 2.2-A14B 快 120 倍。

原始内容

Emad (@EMostaque) 转发了 Enze Xie (@xieenze_jr) 的帖子: 🚀 SANA-Video 2.0 is here! A full-stack optimized video model designed for efficiency — while still delivering high quality. We combine a hybrid architecture closely related to the recent Kimi K3 design, adopt Self-Flow from FLUX 3, and further accelerate it with our own Sol-Engine. Key technical ingredients: 🧠 Hybrid Linear–Softmax Attention 3 gated linear layers + 1 gated-softmax anchor (75% linear / 25% softmax) 🧱 Block Attention Residuals (AttnRes) Boosts deep-layer effective rank by ~12% 🏗️ Unified 5B & 14B models Trained from scratch under very limited resources: • 5B → only 16 nodes of H100 • 14B → 48 nodes of B200 Results 📊 • 84.30 VBench Total • 3.2× faster DiT forward than matched full-softmax at 720p/60s • 720p/5s in 13.06s on a single H100 with Sol-Engine • 120× faster than Wan 2.2-A14B under the same one-H100 setup SANA-Video 2.0 shows that high-quality 720p video generation can be both efficient and practical on a single GPU. 🎬 Project: https://t.co/40uvReE9oE 📄 Paper: https://t.co/cmJZhU64pp 💻 Code: https://t.co/3UlHHqdv0Y Proud of the team! 🎉 More details below 🧵 https://video.twimg.com/amplify_video/2080515508412256256/vid/avc1/1280x720/JRNaonM3vhC0NEeX.mp4?tag=29

相关动态

02

Offloop多智能体超越Claude Code

Offloop 4 人团队公布其多智能体 harness 在 GDPval 基准上以 84.9 分、单任务成本 $1.65 超越 Claude Code(Opus 4.8 得 82.4 分、$14.38) 与 Codex(GPT 5.6 Sol 得 83.3 分、$5.20),并声称在 GDP.pdf(44 分) 与 JobBench(6...

twitter关注列表2026-07-23#评测#技术突破#行业动态
观察
03

Ant Ling 发布 Ling-3.0-flash 模型

Ant Ling 发布 Ling-3.0-flash 混合推理 MoE 模型,124B 参数,仅 5.1B 活跃参数,采用 KDA+MLA 混合注意力机制,支持 256K 上下文。以 1/8 总参数量和 1/12 活跃参数量,在多数基准上匹配或超越其 1T 旗舰模型。SGLang 正与其团队合作提供 day-0 支持。

twitter关注列表2026-07-23#模型发布#大模型#技术突破
观察
05

FLUX 3 发布:真实世界模型

Black Forest Labs 发布 FLUX 3,定位为真实世界模型,将图像、视频、音频、语言整合进多模态 flow 架构,强调模态间的互约束以逼近世界表征。产品线分为 Video(最长 20 秒、原生音频)、Image(即将推出)和 Action(动作预测)。

twitter关注列表2026-07-24#技术突破#模型发布#多模态
观察