Nvidia published a technical deep dive and white paper on July 21 detailing the Olympus core microarchitecture inside its Vera CPU — the processor half of the Vera Rubin platform — along with its first SPEC CPU 2026 results.

What is genuinely new

The headline specs — 88 cores, 176 threads, 164MB of unified L3 and 1.2 TB/s of memory bandwidth — were disclosed back at GTC in March. What landed on July 21 is the internals: a 10-wide decoder fed by 16-instruction fetch and a 48-instruction decode queue, a neural branch predictor handling two branches per cycle, 18 execution pipes (eight integer, six vector/FP, four load, two store), and rename at 10 micro-ops per cycle.

The cache and memory story

Each core gets 64KB of 4-way L1 instruction cache, 96KB of 6-way L1 data cache and 2MB of L2, on top of the shared 164MB L3 and a 3,000-entry L2 TLB. Nvidia also describes a graph prefetcher aimed at pointer-heavy data structures, plus memory renaming and value prediction. The chip is a monolithic die — Nvidia's stated way of avoiding the "chiplet tax" — fed by LPDDR5X it claims delivers 40% lower latency and 5x the bandwidth per watt of DDR5.

The benchmark, read carefully

Nvidia reports a dual-socket Vera at 925 SPECrate2026_int_base against 898 for a dual-socket AMD EPYC 9755 — about 3%. But those numbers are estimates from a pre-production system, self-compiled with GNU 15.2, not official SPEC submissions. Officially submitted dual-socket EPYC systems already exceed 1,000, and Turin Dense parts exceed 1,200.

Where it is going

Early silicon has gone to OpenAI, Anthropic, Oracle Cloud and SpaceX AI; HPE's ProLiant Compute DL394 Gen12 ships in autumn 2026. Nvidia claims up to 1.8x on agentic workloads versus x86.