AI Models

NVIDIA Vera Rubin Boosts Performance Per Watt, Cuts Token Costs

NVIDIA's Vera Rubin platform is ramping up with 300 global partners, delivering 10x more tokens per megawatt than Grace Blackwell NVL72 in benchmarks. CoreWeave, Google Cloud, Microsoft Azure, and Mistral are among early adopters. The platform's extreme co-design across seven chips and five rack trays achieves the highest performance per watt and lowest token cost for AI factories.

Neura News

Neura News

Neura Market Editorial

July 21, 20269 min read
NVIDIA Vera Rubin Boosts Performance Per Watt, Cuts Token Costs

{ "title": "NVIDIA Launches Vera Rubin Platform, Claiming 10x Efficiency Leap in AI Supercomputing", "body": "NVIDIA on Tuesday announced the Vera Rubin platform, a rack-scale AI supercomputer that the company says marks a generational leap in performance and energy efficiency. Production is ramping up with racks already running at CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, and Nebius, and early benchmark results show dramatic gains over the previous generation.\n\nThe Vera Rubin NVL72 system, built from chip to grid, is designed to deliver the highest performance per watt and the lowest token cost, according to NVIDIA. The company claims it has assembled the largest, most mature rack-scale supply chain ever for the platform, spanning more than 350 factory sites across 30 countries.\n\n## A 10x Leap in Efficiency\n\nCoreWeave, an AI cloud provider and NVIDIA partner, was the first to bring up and validate the Vera Rubin NVL72. The company ran a benchmark on the DeepSeek-R1 model and reported a 10x improvement in tokens per second per megawatt compared with the Grace Blackwell NVL72.\n\nThe benchmark underscores a central metric for AI factories: tokens per megawatt. Compared with the NVIDIA GB200 NVL72, the Vera Rubin NVL72 delivers up to 10x more tokens per megawatt and one-tenth the cost per million tokens. This efficiency gain is critical as AI workloads grow, with token consumption rising rapidly across the industry.\n\nGoogle Cloud also announced the A5X instance, powered by the Vera Rubin NVL72 and the Google Virgo Network. The bare-metal instances deliver up to 10x lower inference cost per token and 10x higher token throughput per megawatt than the prior generation. The first customer is Ineffable Intelligence, a London startup developing "superlearner" systems.\n\n"The next era of research requires the next era of hardware," said Lasse Espeholt, cofounder of Ineffable Intelligence. "We feel privileged to work with the teams at NVIDIA and Google Cloud, who were able to grant us early access to Vera Rubin. The support across both teams has been unmatched; we were up and running almost immediately and are already testing infra for our superlearners."\n\nThe A5X instances use NVIDIA ConnectX-9 SuperNICs with next-generation Google Virgo networking. Clusters can scale to tens of thousands of NVIDIA Rubin GPUs within a single site and up to nearly a million GPUs across multisite configurations. This scalability is designed to support the largest AI training and inference workloads.\n\nOracle Cloud Infrastructure is also deploying the Vera Rubin NVL72, with early access customers including AI startups and enterprise clients. Microsoft Azure is integrating the platform into its global data center fleet, with a focus on high-performance AI workloads. Nebius, an AI cloud provider, received the first Vera Rubin NVL72 system at its Finland AI Factory and is deploying the Spectrum-6 102.4T Ethernet switch. Nebius is an early entrant bringing the Vera Rubin platform to Europe and the U.S.\n\n## Extreme Codesign Across Seven Chips\n\nThe Vera Rubin platform is the result of what NVIDIA calls extreme codesign, integrating seven chips and five rack trays as a single system. The components include the Vera Rubin NVL72, the Vera CPU rack, the Groq 3 LPX, the Spectrum-6 SPX, and the Vera BlueField-4 STX. This integration allows the system to function as a unified supercomputer rather than a collection of discrete parts.\n\nThe NVIDIA Vera CPU uses a custom Olympus core that delivers 2x single-threaded performance, 3x core-to-core bandwidth, and 40% lower memory latency versus competing chiplet designs. This makes it suitable for both AI orchestration and traditional HPC workloads.\n\nThe sixth-generation NVLink scale-up fabric provides more than 2x throughput on complex workloads, 3x lower latency, and 10x higher packet rates than off-the-shelf Ethernet. This fabric is critical for connecting GPUs within a rack, enabling the 260 TB/s all-to-all NVLink 6 fabric that allows the rack to behave as a single unified accelerator.\n\nThe Spectrum-X Ethernet, built on 102.4T Spectrum-6 switch systems and 1.6T ConnectX-9 SuperNICs, enables 1.6x higher RDMA bandwidth than off-the-shelf Ethernet. It includes adaptive routing, advanced congestion control, telemetry, and open OS support. This networking stack is designed for AI factories that require low-latency, high-bandwidth communication across thousands of nodes.\n\nNVIDIA Photonics with co-packaged optics for scale-out is the industry's first such switch in volume manufacturing. It adds 5x lower power and 10x higher mean time between interruptions (MTBI) versus pluggable transceivers. This technology reduces energy consumption and improves reliability in large-scale deployments.\n\nSpectrum-XGS Ethernet delivers 1.9x multi-site throughput, enabling efficient scaling across data centers. NVLink Fusion opens the NVIDIA infrastructure platform to third-party XPUs, allowing customers to mix and match accelerators from different vendors. This openness is intended to give customers flexibility in building their AI infrastructure.\n\nSpaceXAI and Tesla are among the first to bring in Spectrum-6 switches, indicating adoption beyond traditional cloud providers. Lambda is an early adopter of NVIDIA Photonics, using the technology to improve network efficiency in its cloud services.\n\n## Engineering for Scale and Efficiency\n\nThe Vera Rubin NVL72 system has no cables, fans, or hoses in the tray. Compute tray assembly time has been cut from hours to one minute. This simplification reduces manufacturing complexity and improves reliability by eliminating moving parts and cable connections that can fail.\n\nThe system is designed for a 45-degree Celsius liquid cooling inlet temperature, enabling chiller-free dry-cooler operation. Higher-temperature dry cooling, combined with closed-loop liquid cooling, saves millions of gallons of water per megawatt annually for new AI factories. This is a significant environmental benefit as AI data centers face scrutiny over water usage.\n\nThe Vera Rubin NVL72's 260 TB/s all-to-all NVLink 6 fabric enables the rack to behave as a single unified accelerator. This design eliminates the need for complex interconnects between GPUs within a rack, reducing latency and improving performance.\n\nCoreWeave is deploying the Spectrum-X Ethernet SN6600-LD as the switching fabric for its Vera Rubin NVL72 systems. Built on the 102.4 Tb/s Spectrum-6 switch chip and liquid-cooled, it delivers 1.64 Pb/s per rack with 100% more capacity than previous-generation air-cooled switches. This allows CoreWeave to scale its AI cloud services efficiently.\n\nThe supply chain for the Vera Rubin platform spans more than 350 factory sites across 30 countries, according to NVIDIA. This global network is designed to meet the high demand for AI infrastructure, with production already ramping up to support multiple cloud providers and enterprise customers.\n\n## Agentic AI and the Vera CPU\n\nDeepInfra, a cloud platform for AI inference and an early access participant in NVIDIA's open AI ecosystem, processed nearly 5 trillion tokens a week, with about 30% driven by agentic systems. Agentic systems can consume up to 15x more tokens than traditional AI applications. This trend is driving demand for more efficient inference hardware.\n\nDeepInfra ran benchmarks on the NVIDIA Vera CPU. The results showed that the Vera CPU supports up to 1.6x more concurrent AI agents at the same quality of service and up to 2.2x faster orchestration than alternative CPUs. This makes it suitable for running complex AI workflows that require both high throughput and low latency.\n\nThe Vera CPU is designed to handle the orchestration layer of AI systems, managing tasks such as model routing, data preprocessing, and agent coordination. By offloading these tasks to a dedicated CPU, the system can free up GPU resources for inference and training.\n\n## European AI Infrastructure Expansion\n\nThe Vera Rubin platform is the foundation for an expanded partnership between Microsoft and Mistral for European AI infrastructure. A new multibillion-dollar agreement is focused on expanding AI infrastructure in Europe. Mistral is adding GPU capacity drawing on thousands of the latest NVIDIA Vera Rubin GPUs.\n\nMistral Medium 3.5 and OCR 4 are now available in Microsoft Foundry, and Mistral models are integrated into Microsoft Copilot Studio. This integration allows European enterprises to access advanced AI models running on Vera Rubin hardware.\n\nNebius, an AI cloud provider, received the first Vera Rubin NVL72 system at its Finland AI Factory and is deploying the Spectrum-6 102.4T Ethernet switch. Nebius is an early entrant bringing the Vera Rubin platform to Europe and the U.S., positioning itself as a key player in the European AI cloud market.\n\nSpaceXAI and Tesla are among the first to bring in Spectrum-6 switches, indicating adoption beyond traditional cloud providers. Lambda is an early adopter of NVIDIA Photonics, using the technology to improve network efficiency in its cloud services.\n\nNVIDIA GTC Berlin registration is open for October 20-22, where the company is expected to showcase the Vera Rubin platform and other AI technologies. NVIDIA Jetson was announced on July 28, 2026, expanding the company's AI portfolio into edge computing.\n\n## Related on Neura Market\n\n- NVIDIA Vera Rubin Platform\n- AI Infrastructure and Cloud Computing\n- European AI Ecosystem" }

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

More from Neura News

Product Launch

Acer Unveils Veriton RI110 Mini Workstation for Local Agentic AI

Acer unveiled the Veriton RI110 AI Mini Workstation on September 2, 2026, in Berlin. This compact desktop, featuring an Intel Core Ultra X7 processor and Intel Arc B390 graphics, supports local inference of AI models up to 120 billion parameters. It is designed for hybrid agentic AI workloads, combining local processing with cloud resources, and includes the Qubi Claw software suite for secure, autonomous AI tasks. The system offers up to 96 GB of LPDDR5X memory, 4 TB of SSD storage, and extensive connectivity options including OCuLink, Wi-Fi 7, and dual LAN ports. Availability begins in North America in Q4 2026 and EMEA in Q1 2027.

Sep 2·4 min read