Signal Brief

推荐最高密度GPU编程视频

swyx 转发 levi 的帖,强烈推荐 Ben Spector 的 4.5 小时 CUDA+ThunderKittens 教学视频,称其是最高密度的 GPU 编程讲座,目前仅 5k 观看,被严重低估。

twitter关注列表 swyx (@swyx) 发布 2026-07-05 收录 2026-07-06 观察

一句话判断

该视频结合顶尖研究者的实战经验,内容密度极高,值得正在学习 GPU 编程的从业者观看原文。

核心信息

swyx 转发 levi 的帖,强烈推荐 Ben Spector 的 4.5 小时 CUDA+ThunderKittens 教学视频,称其是最高密度的 GPU 编程讲座,目前仅 5k 观看,被严重低估。

原始内容

swyx (@swyx) 转发了 levi (@levidiamode) 的帖子: 183/365 of GPU Programming This 4.5 hour lesson on CUDA + ThunderKittens by @bfspector (TK co-author, Stanford PhD student) is one of the best educational videos on kernels out there (think @karpathy level video but solely focused on GPU stuff). It's so criminally underrated, I cannot believe this video only has 5k views. It's the highest density GPU programming lecture I've come across in my ~6 months of learning. For better or for worse, Ben talks (and thinks) at the speed of light. But thanks to @qamcintyre it's more of a dialogue between Ben who introduces concepts/ideas based on his experience creating TK and Quinn who intervenes from time to time with pointed questions that help ground the discussion in fundamentals. Regardless of whether you're just starting out with writing kernels or have some experience, it's an incredible resource because it gives you a behind-the scenes-look of someone who's spent years thinking about how to maximize what you can get out of the hardware (read his Hazy blog post "We Bought the Whole GPU, So We're Damn Well Going to Use the Whole GPU" for context). Even though I was sort of familiar with the topics of the lesson (e.g. GPU architecture and memory consistency), I still came away with pages of notes and questions around TMA accelerated pipelining, coscheduling instructions, pre vs post Hopper, levels of virtualization, cache lines, barriers, TK primitives, et cetera. As @suryaasub so aptly put it in the YT comments: "greatest video of all time". ![photo](https://pbs.twimg.com/media/HMd3yXSXYAAm1r3.jpg) > **引用原帖 levi (@levidiamode):** > 182/365 of GPU Programming > Preparing my first submissions to the eigendecomp leaderboard today, and gotta say this GPU Mode challenge feels different already. The speed at which competitors have cleared sub 20ms and now 10ms even though we're still more than a week out is quite remarkable (and a step up from the previous QR challenge). Wonder how much of that has to do with Fable coming back a couple days ago. > Don't have any remaining access to Fable (burned through my small usage limits on a couple of side projects) but will try to get the most out of my harness with GPT 5.5 and Opus 4.8. This second challenge is mainly a way for me to iterate on the harness and learn more about linear algebra as well as Blackwell kernels. > https://x.com/levidiamode/status/2073556115170656613

相关动态

02

Kimi K3非对称竞争分析

Kimi K3在Arena前端/WebDev人类盲测中以1679 Elo排名第一,领先Fable 5(1631)和GPT-5.6 Sol(1618),但在综合智能指数中仅排第四(57.1分),落后Fable 5(59.9)和GPT-5.6 Sol(58.9)。K3定价15美元/百万输出tokens,远低于Fable 5的50美元,并计划7...

twitter关注列表2026-07-18#AI#模型#技术突破
观察
04

BestBlogs 早报 · 07-18

月之暗面发布Kimi K3,2.8万亿参数,896选16的Stable LatentMoE,上下文100万token,接近Fable-5但仍落后最强闭源模型,完整权重7月27日前开源;VentureBeat调查显示54%企业已发生AI代理安全事件;xAI开源Grok Build(84万行Rust代码)并残留上传用户代码痕迹;Cursor评...

twitter关注列表2026-07-17#AI#模型发布#开源
观察
05

Grok 4.5成本效率对比评测

Artificial Analysis数据显示,Grok 4.5每次任务成本仅0.31美元,而Claude Fable 5为2.75美元、Claude Opus 4.8为1.80美元、GPT-5.6 Sol为1.04美元、Kimi K3为0.95美元,Grok 4.5成本分别低约9倍、6倍、3倍和3倍;在FrontierSWE上,Grok...

twitter关注列表2026-07-17#AI#模型#评测
观察