GREAT WORK BY @GPU_MODE 🚨 FOR LAUNCHING THE $1.1mil AMD KERNEL HACKATHON.
The GPUMODE Readonflow Team’s kernels improved end-to-end MI355X performance by over 2x. We explain the optimizations below. 1/4🧵
@SemiAnalysisGREAT WORK BY @GPU_MODE 🚨 FOR LAUNCHING THE $1.1mil AMD KERNEL HACKATHON.
The GPUMODE Readonflow Team’s kernels improved end-to-end MI355X performance by over 2x. We explain the optimizations below. 1/4🧵
The Wild Wild West Of LEGO Datacenters
Everyone Says They're Modular,Do The Vendor Claims Hold Up?Zuck's Tents, AWS's Houdini,60GW+ Modular Capacity Tracked,Full Vendor Landscape Mapping,Vertiv's 2x Content Uplift Per MWAMD's warrant deals with OpenAI and Meta usually get described as equity sweeteners. Run the math, and these look more like a rebate that matches the price of the compute itself, with up to a 105% discount for OpenAI. (1/3)🧵

Vera Rubin NVL72 vs GB200 NVL72? Inference TCO & Architecture Analysis:
Rubin LUT Based Tensor Core, Feynman,Rack Scale, Perf Per MegaWatt, Perf Per Dollar,Software Improvements, Public Rubin Software,PyTorch, vLLM, OpenAI Triton Read Now:ALERT🚨🚨: META's CUSTOM AMD MI400-series chip will be half the size of a normal MI455X chip. It is "optimized" for recsys workloads and $/Memory Bandwidth. It will use ~144GB of HBM instead of 432GB. We break it down below👇️ 1/7🧵

Datacenters have always kept reciprocating engines on site as emergency standby, kicking in when the grid drops and running a handful of hours a year. That role is changing, as recips are increasingly being repurposed to prime power, energizing datacenters around the clock. We estimate that recip OEMs (e.g., Caterpillar, INNIO, Cummins) have been contracted to supply ~1 GW of BTM power this year, and 4+ GW in each of 2027 and 2028. Our SemiAnalysis Energy Model tracks these BTM contracts OEM-by-OEM, project-by-project. Yet, that is dwarfed by the opportunity to come. (1/3)🧵

Great work by the AMD @sgl_project team on enabling nightly disaggregated serving CI to improve code quality! It has already caught and prevented 2 massive bugs from reaching customers, as we explained before 👇️ 1/7🧵

Similar to the panic over DeepSeek R1, some uneducated people think Kimi K3’s use of linear attention (KDA) is bad for NVIDIA, HBM, DRAM, and networking because it has relatively lower KV-cache requirements. The opposite is true, and we explain why below. 👇️ 1/8🧵

BREAKING: Broadcom is diversifying away from long-time foundry partner TSMC. Lego will be the new manufacturing partner for Broadcom's flagship hyperscaler ASIC products in 2028. Broadcom President of Semiconductor Solutions, Charlie Kawwas, has already showed off package samples at RAISE Summit in Paris last week. With TSMC supply constrained, customers are alternative sources of chip supply from unexpected places. One of the major differentiators Lego brings is industry leading defect repair, "plug and play" chiplet interoperability, as well as built in self-alignment technology for 3D stacking. Mattel and Hasbro were also under consideration, but Lego's proven track record won out in the end. We have ordered multiple Millennium Falcon Sets set to our Portland teardown lab for competitive teardown analysis.