一句话判断
建议关注H800芯片在训练和推理中的实际性能优化研究进展,验证多芯片并行训练的稳定性
核心信息
该分析指出,训练数据量可能比早期阶段的H800芯片更大又更稳定(若原始为15T则扩展至30-45T令牌量)。同时指出多芯片优化(如暴露局部芯片解决方案、64芯片优化),采用MXFP4精度、静态形状和无主机端协同等架构设计,进一步强调了训练效率的跃升:数千块芯片就足以完成前沿模型训练,推动了对RAM需求的影响分析(910cs节点配备1TB RAM但可能实现低于400GB包装)。
原始内容
Of course total data etc would be a multiple of that but this is likely on the same H800s they trained K2.x on but larger & more stable
Multiple of that as tokens scale (15tr K2 => 30-45tr potentially)
We also see local chips being used on kernel for "supernode" deployment in their blog with 64 chip optimisation (Huawei? Could also be Alibaba Panjiu) & why MXFP4, static shape/no host synchronisation etc
https://t.co/XCWYynp4ld
1e25 flop level training runs being able to build frontier models is not what is expected/the common story - this is doable on a few thousand chips not the 10k-100k-1m chips we hear about for training runs.
The other thing underappreciated with the architectural things are how optimised the inference is likely to get with the impressive quantisation we are seeing, this can get it down to a < 400Gb RAM package potentially (910cs have 1tb of RAM per node)
https://t.co/Jo7QpA01yq
> **引用原帖 Emad (@EMostaque):**
> This is v bearish for RAM companies.
> We saw strong ternary numbers from @PrismML today (27b dense on a mobile!) but the tiny degradation (~5%) from 16 bit to 1 bit precision by Tencent on a near frontier 300b model is the biggest news of the day.
> Star https://t.co/6xqbYh7Mez
> https://x.com/EMostaque/status/2077103737953304683