NVIDIA has officially introduced the Rubin platform, a next-generation AI supercomputer architecture that integrates six specialized chips, including the new Vera CPU and Rubin GPU. Designed through extreme hardware–software co-design, Rubin delivers massive gains in efficiency, offering up to 10× lower cost per inference token and requiring 4× fewer GPUs for training mixture-of-experts models compared to the previous Blackwell platform.
The Rubin platform combines advanced components such as NVLink 6 switches, ConnectX-9 SuperNICs, BlueField-4 DPUs, and Spectrum-6 Ethernet switches to create a powerful rack-scale AI system. It also introduces next-generation Transformer Engine technology, third-generation confidential computing, and contextual AI storage, all aimed at accelerating agentic AI reasoning and large-scale model workloads.
NVIDIA plans to ship Rubin-based systems in configurations like the Vera Rubin NVL72 rack and HGX Rubin NVL8, enabling scalable deployment for both AI training and inference. These systems are built to support the growing demands of frontier AI models and are expected to be available through NVIDIA’s partners starting in the second half of 2026.
Major cloud providers and AI research organizations—including leading hyperscalers and AI labs—are expected to adopt the Rubin platform to power their next wave of AI innovation. With Rubin, NVIDIA reinforces its position at the center of global AI infrastructure, setting a new benchmark for performance, efficiency, and scalability in AI supercomputing.
Source: Nvidia Newsroom
