AMD launched Helios, its first rack-scale AI system, at Advancing AI 2026 in San Francisco on Thursday. Lisa Su said the rack is in production now, with partner shipments starting at the end of Q3 and ramping through Q4 into the first half of 2027.
What's in the box
Seventy-two Instinct MI455X accelerators and eighteen sixth-generation EPYC "Venice" CPUs across 18 compute trays. Total memory is 31TB of HBM4 at 1.7 PB/s aggregate bandwidth, for 2.9 exaflops of peak MXFP4 and about 1.4 exaflops at FP8. Scale-up runs at 260 TB/s over UALink-over-Ethernet through twelve Broadcom Tomahawk 6 switches; scale-out adds 43 TB/s. It is an OCP Open Rack Wide double-wide unit weighing roughly 7,000 pounds and drawing 225-245 kilowatts under load on a 50-volt liquid-cooled DC bus.
The parts
Each MI455X carries 432GB of HBM4 across twelve stacks at 23.3 TB/s, with 320 billion transistors in twelve chiplets — eight compute dies on TSMC N2 plus four on N3. AMD rates it up to 4x the MI355X on MXFP4. Venice tops out at 256 cores and 512 threads on TSMC's 2nm process, making it the first x86 server CPU in volume production on that node, with roughly 70% more compute than Turin. Networking is Pensando: three Vulcano 800G NICs per GPU, and a Salina 400 DPU per blade.
The claim against Nvidia
Against Vera Rubin NVL72, AMD's own modeling claims 15% more AI compute, 50% more HBM capacity, 50% more scale-out bandwidth per GPU and up to 30% more inference tokens per dollar — hitting the same 3.6 TB/s scale-up per GPU with a third as many switch ASICs. Analysts at Futurum estimate a Helios rack at $5-5.5 million against $3.5-4 million for Rubin; AMD published no price.
The gap
AMD holds roughly 4.5% of the data-center GPU market against Nvidia's 95%-plus. Named customers include OpenAI, Anthropic, Meta, Microsoft, Oracle, HUMAIN, TensorWave, Vultr and Cirrascale. IDC's Ashish Nadkarni said AMD is "making steady progress to be the strong No. 2." Helios 500, with MI500 and EPYC "Verano," is slated for 2027.
