Nvidia’s Backstop Universe
Heads I Win, Tails Who Loses?The $11T AI Buildout,Nvidia’s Backstop Economics,and the Limits of Nvidia’s Balance Sheet

Nvidia’s Backstop Universe
Heads I Win, Tails Who Loses?The $11T AI Buildout,Nvidia’s Backstop Economics,and the Limits of Nvidia’s Balance SheetSpec Decoding (DSpark) finally works with Pipeline Parallelism in vLLM!! 🔥 A lot of the GPU poor/proletarian have been asking for this feature for a while now, but it took the recent Kimi K3's massive 2.8T parameters affecting the GPU middle class (B200) for it to be implemented!
Before this change, GPU poors that needed to use pipeline parallelism, as their weights didn't fully fit on 1 server, would not be able to take advantage of DSpark.Shoutout to yongqinwang-cmd & Inferact for implementing it in vLLM!
TPU Inference Externalization Full Steam Ahead
InferenceX, Up to 50% Better Performance per DollarRapid Externalization of TPU stack, Growing Customer BaseIronwood, TPUv8i, Reducing CUDA MoatNVIDIA LPU supports 3 types of disaggregated inferencing:
1. Rubin Prefill + LPU Decode for the fastest interactivity2. Rubin Prefill + Rubin Decode Attention + LPU Decode FFN for the middle of the curve3. Rubin Prefill + Rubin Decode Verification + LPU Drafter for the middle-left of the curveFor low interactivity, raw Rubin still takes the win. Looking forward to seeing Rubin + LPU performance curves on open-source agentic benchmarks like AgentX.BREAKING NEWS: AMD middle managers are wasting over 20+ engineers’ time trying to grow their fiefdoms & get promotions on a useless side project called their “ATOM Inference Engine,” instead of spending that effort on inference engines that actual customers like Meta/xAI actually use, such as SGLang/vLLM.
Dozens of AMD engineers, including principal engineers and AMD Fellows, have reached out to us privately about how, due to office politics and backing from a subset of middle managers at AMD, they are prevented from talking internally about their concerns over middle management wasting 20+ full-time engineers’ time on a side project with basically no customers. The engineers on this side project are 10x engineers and instead should be working on customer engines like vLLM/SGLang.The middle managers will try to cope by referencing the singular customer that uses it in prod, Alibaba on Qwen, but when digging deeper, it is just a side-business enterprise unit deploying ATOM in a limited capacity instead of the main Qwen business unit.It is incredibly disrespectful to AMD engineers for AMD management to pull resources & limited internal GPU R&D dev clusters away from engineers supporting customers on production engines like vLLM/SGLang to spend them on growing their fiefdoms & playing office politics.We have written extensively about how forecasted AI IT and AI Datacenter capex will drive a record breaking amount of funding needs and an unprecedented growth in debt issuance.
By the end of this year, cumulative lifetime AI Capex will have reached nearly $3T USD and we expect this number to grow to $11T+ by the late 2020s, with total outstanding AI Debt reaching $7T+. 2026 will be the first year in which AI Debt Financing becomes the second largest asset related debt market, after only the US Residential Mortgage Market. It is also the last year in which total AI related debt issuance will be less than 10% of total US Fixed Income Issuance (this includes state and federal bonds, MBS/RMBS, corporate bonds, agency, muni and ABS). (1/3)🧵Great to see @dotsstudioai from the popular Chinese social media app (@xiaohongshu, known as RedNote in English) join the open-weight community! ❤️ The team is still trying it out for day-to-day use to see the actual quality of the model or if it is eval maxed.
Regardless, it is always great to see another company join the open community!How do different operating systems and microarchitectures affect the end user experience with agentic AI?
We looked through 2.91M actual Claude tool use responses to find out.By our data, it appears that Linux is typically faster than Windows by a factor of about 3x in execution of shell commands. (1/2)🧵NPO presents an interim solution under the transition from pluggable to true CPO. NPO has certain benefits over CPO that bypass current production and reliability challenge of CPO, while maintaining most of the benefits CPO provide.
Pros:🟠 Better serviceability (field replacable)🟠 Lower blast radius (blast radius limited to the individual NPO module)🟠 Easier assembly (Optical Engine are packaged separately from the Switch ASIC/XPU)Let's look at how NPO differentiate from CPO architecturally. (1/3)🧵THIS IS AN ENGAGEMENT BAIT & HOT TAKE 🚨Percy Jackson is a better movie than Odyssey in terms of storytelling. Odyssey is one of Christopher Nolan's worst films. The storytelling was so poor that they had to distract viewers by adding IMAX 70mm cameras.
Also FYI for those of y’all who did not have the luxury of watching in IMAX 70mm film, you are missing 40% of the movie.Most of AMD’s top 10x AI engineers are in Shanghai 🔥 This includes AMD’s MoRI collective and UMBP KV-cache offloading and pooling team, its disaggregated-applications forward-deployed engineering team, and other AMD teams that understand how to approach inference engineering from first principles.
Many of the most important components of the ROCm software stack are built in mainland China.The datacenter fight in the US has so far run through county boards and city councils. But that ended on July 14, when Gov Kathy Hochul signed Executive Order 62 and New York became the first state to implement a datacenter moratorium. (1/5)🧵
SemiAnalysis Datacenter team tracks all dc sites. Datacenter team started noticing new builds missing the gensets. Datacenter team checked the Industrials Model and now tracks 15GW+ of capacity with no gensets, no central UPS, or neither. (1/3)🧵
How often does Claude use tools?
We looked through 2.27M Claude responses we collected to find out.The Opus models show a clear downward trend in tool calls from Opus 4.6 to Opus 4.8. However, Fable 5 stands out against that pattern, averaging 1.00 tool calls per response compared with 0.76 for Opus 4.8 and 0.79 for Opus 5. (1/2)🧵We believe the MI455 is the first known chip to ship with active Local Silicon Interconnect (LSI) bridges, the chiplet-to-chiplet links inside a CoWoS-L package.
In CoWoS-L's short history, those bridges have always been passive: wiring and capacitors, nothing more. Active LSIs carry real circuitry. At ISSCC 2026, TSMC demoed active bridges with a low-power repeater that regenerates signals mid-channel.Why it matters: once the bridge shares the signal-integrity burden, the PHYs on the top dies can shrink meaningfully at near-zero energy cost, reclaiming leading-edge silicon and shoreline for compute and memory.The giveaway that this is a shipping product rather than a research vehicle: the interposer TSMC showed matches the MI455 exactly. Two base dies, twelve HBM4 stacks, two IO dies.AMD MI355X vLLM HAS BEATEN B200 vLLM ON KIMI K2.5 (THE SAME MODEL ARCH AS XAI CURSOR COMPOSER 2.5)🚀🚨
This uses upstream AMD kernels from the @GPU_MODE community. We explain below 👇 1/6🧵Indium Phosphide (InP) LASERs are suddenly catching everyone's interest since it suddenly became one of the more strategic components in computing. As optical connectivity moves ever closer to the package, every optical engine and laser source inevitably runs on InP LASERs. But supply isn't keeping up with demand, particularly for Continuous Wave (CW) Distributed Feedback (DFB) LASERs which enable near-package optical connectivity. (1/2)🧵
Ever wonder why PCB folks throw around "M8" or "M9" like a spec sheet magic word?
The "M" traces back to Panasonic's Megtron series — long the benchmark for low-loss, high-speed copper clad laminate (CCL).Each generation from Megtron 4 to 8 progressively lowered the dissipation factor (Df), meaning less signal loss at high frequencies.Because Megtron set the bar, it became industry shorthand for a performance tier rather than a Panasonic-exclusive part number. Say "M8 grade" today, and it simply means "meets Megtron 8 electrical specs" — whether it's made by Panasonic, EMC, or Doosan.As AI servers demand ever-higher SerDes speeds, high-tier CCL has become the critical bottleneck material. That makes it one of the tightest, fastest-growing segments in the entire hardware supply chain.GREAT WORK BY @GPU_MODE 🚨 FOR LAUNCHING THE $1.1mil AMD KERNEL HACKATHON.
The GPUMODE Readonflow Team’s kernels improved end-to-end MI355X performance by over 2x. We explain the optimizations below. 1/4🧵The Wild Wild West Of LEGO Datacenters
Everyone Says They're Modular,Do The Vendor Claims Hold Up?Zuck's Tents, AWS's Houdini,60GW+ Modular Capacity Tracked,Full Vendor Landscape Mapping,Vertiv's 2x Content Uplift Per MWAMD's warrant deals with OpenAI and Meta usually get described as equity sweeteners. Run the math, and these look more like a rebate that matches the price of the compute itself, with up to a 105% discount for OpenAI. (1/3)🧵
Vera Rubin NVL72 vs GB200 NVL72? Inference TCO & Architecture Analysis:
Rubin LUT Based Tensor Core, Feynman,Rack Scale, Perf Per MegaWatt, Perf Per Dollar,Software Improvements, Public Rubin Software,PyTorch, vLLM, OpenAI Triton Read Now:ALERT🚨🚨: META's CUSTOM AMD MI400-series chip will be half the size of a normal MI455X chip. It is "optimized" for recsys workloads and $/Memory Bandwidth. It will use ~144GB of HBM instead of 432GB. We break it down below👇️ 1/7🧵
Datacenters have always kept reciprocating engines on site as emergency standby, kicking in when the grid drops and running a handful of hours a year. That role is changing, as recips are increasingly being repurposed to prime power, energizing datacenters around the clock. We estimate that recip OEMs (e.g., Caterpillar, INNIO, Cummins) have been contracted to supply ~1 GW of BTM power this year, and 4+ GW in each of 2027 and 2028. Our SemiAnalysis Energy Model tracks these BTM contracts OEM-by-OEM, project-by-project. Yet, that is dwarfed by the opportunity to come. (1/3)🧵
Great work by the AMD @sgl_project team on enabling nightly disaggregated serving CI to improve code quality! It has already caught and prevented 2 massive bugs from reaching customers, as we explained before 👇️ 1/7🧵