Nvidia gave the first architectural disclosure of the Rubin GPU at Hot Chips on Monday evening: two reticle-sized dies on TSMC 3nm, 336 billion transistors, roughly 1.6x Blackwell, carrying up to 288 GB of HBM4 at 22 TB/s — about 2.8x Blackwell's HBM3e bandwidth.
The rack, and the building
At rack level, NVLink 6 moves 3.6 TB/s all-to-all per GPU with 130 TFLOPS of in-network compute, on 800 VDC distribution and a 45°C liquid inlet. Nvidia claims 40% more GPUs per provisioned watt through power smoothing. Then there is a second tier of numbers — 2 ZFLOPS of NVFP4 inference, 1.4 ZFLOPS of training, 11 PB of HBM4, 800 PB/s — and those are stated at 100 MW AI-factory scale.
What the conventional framing gets wrong
Three errors, and they compound. First, 11 PB of HBM4 is a building, not a rack and not a system; it is the memory installed across a 100-megawatt facility. The 2 ZFLOPS is the same denominator. Both will be reprinted as rack figures within the week. Second — and this is the one that inverts the story — 288 GB is not more memory than you can buy today. It is the same capacity as Blackwell Ultra. Anyone reading "HBM4" as "bigger models now fit on one GPU" has it backwards; what improved is how fast weights move, not how many of them sit there. Third, the "up to 30x" figure is not a speedup. It is tokens per megawatt at high interactivity, measured on one vendor-chosen agentic workload — DeepSeek-v4-PRO at 140K-plus context — and 10x is the same curve at a different point on it.
The metric is the tell
Nvidia has stopped leading with FLOPS and started leading with tokens per watt at a given latency. That is not marketing drift. It is an accurate read of what now binds its customers: not the price of silicon, but how much power a site can get connected, and how many tokens can be extracted from each megawatt once it is. A vendor that quotes throughput per watt is telling you which constraint it thinks you are actually shopping against.
One thing this is not
Vera Rubin was already declared in production at Computex earlier this year. This is a technical disclosure about a shipping product, not a preview of an unreleased one — which makes the factory-scale framing of the headline numbers a choice rather than a necessity.
