NVIDIA Vera Rubin is a next-generation AI compute platform, the successor to the Grace Blackwell architecture. It pairs the new Rubin GPU (3 nm process, HBM4 memory, about 50 petaFLOPS of FP4 per GPU) with the new, custom Vera CPU, which features 88 NVIDIA Olympus (Arm) cores, an LPDDR5X memory subsystem and the second-generation NVIDIA Scalable Coherent Fabric. Vera is the successor to the Grace CPU.
The flagship rack-scale configuration, Vera Rubin NVL144, combines 144 Rubin GPUs (72 packages) and 36 Vera CPUs in a single rack. The system delivers up to 3.6 exaFLOPS of FP4 for inference and 1.2 exaFLOPS of FP8 for training, roughly 3.3x the performance of the current GB300 NVL72 platform. It provides HBM4 memory with about 13 TB/s of bandwidth and around 75 TB of fast system memory.
Connectivity is provided by NVLink 6 (about 260 TB/s of rack-scale scale-up bandwidth), NVIDIA ConnectX-9 NICs (about 28.8 TB/s) and the BlueField/Quantum-X platform. It runs on the NVIDIA software stack (CUDA, AI libraries). Rubin and Vera Rubin NVL144 are expected in the second half of 2026; the next roadmap step is Rubin Ultra NVL576 (2027, about 100 PFLOPS FP4 per GPU).

Rack-scale AI system
Which group NVIDIA Vera Rubin (NVL144) belongs to and how it is built
A component type covering complete rack-scale AI systems: dozens of GPUs/accelerators tied to CPUs and high-speed networking (scale-up/scale-out) within a single rack, designed for frontier-AI training and inference.
A rack-scale design in which GPUs (AI accelerators), CPUs and NICs are co-designed and densely interconnected (e.g. via UALink, Infinity Fabric, Ethernet), typically in an open rack format (e.g. double-wide ORW), to maximize bandwidth, data movement and scalability for frontier AI.
Basic physical properties of NVIDIA Vera Rubin (NVL144) — dimensions, weight and materials
Other hardware parts related to NVIDIA Vera Rubin (NVL144)

Rack-scale, liquid-cooled computing system built around 72 NVIDIA Blackwell Ultra (B300) GPUs and 36 NVIDIA Grace CPUs, designed for large AI model inference and training.

Rack-scale, liquid-cooled computing system built around 72 NVIDIA Blackwell (B200) GPUs and 36 NVIDIA Grace CPUs, designed for training and inference of large AI models.