Nvidia is promoting its new Vera Rubin chip system with performance benchmarks this week, positioning itself as a supplier of complete AI data center systems, just days before rival AMD holds its annual product event in San Francisco on Thursday.
The push comes as Nvidia executives tout increased power and efficiency for the system, which succeeds the Grace Blackwell hybrid superchip. The company held a lengthy technical workshop last week at its headquarters in Santa Clara, California, for a small group of journalists.
The Shift to CPUs for Agentic AI
GPUs remain the main hardware for training and running AI models. But the shift toward agentic systems, where AI performs multi-step tasks, has increased demand for CPUs. Nvidia is positioning itself as a supplier of CPUs for AI agents, a move that puts it in more direct competition with established CPU makers.
Vera Rubin is designed to offer one CPU for every two GPUs. In a single Vera Rubin NVL72 system, there are 36 Vera CPUs for every 72 Rubin GPUs. Nvidia is also selling the Vera CPU as a stand-alone product, and reportedly told Chinese customers the chips could be ready as soon as August.
The briefings were led by Ian Buck, Nvidia's vice president of accelerated computing and architect of CUDA software. Meetings were held in CEO Jensen Huang's executive briefing center, where desks nearby were piled with bags of Taiwanese snacks from Computex, the annual semiconductor trade show in Taipei. Huang did not appear at the workshop. He was in Japan announcing partnerships with Japanese firms for AI robotics.
Buck struck a confident tone about the company's roadmap. "We're on a road map to crank out new architectures, not just GPUs but CPUs," he said. He added, "We're going to keep innovating, because it's do this or die in Silicon Valley."
Performance Claims and Benchmarks
Nvidia claims the Vera Rubin NVL72 will process 10 times as many tokens per watt as Grace Blackwell. The company also says the Vera CPU is faster at agentic AI tasks compared to AMD and Intel CPUs, though the benchmarks used slightly older generations of competitors' chips.
Localized memory subsystems in Vera Rubin offer nearly three times as much memory bandwidth as Blackwell. This matters amid an ongoing shortage of high-bandwidth memory, a constraint that has affected the broader chip industry.
The system was unveiled in the spring of 2025. Huang has said Vera Rubin is ramping to "full production" and will ship in the second half of this year. Early customers include Microsoft, OpenAI, and Oracle. OpenAI already has one Vera Rubin rack in use.
Nvidia is sensitive to delays after Blackwell chips reportedly overheated in customized server racks. Those overheating issues forced design changes and shipment pushbacks, and the company appears eager to avoid a repeat with Vera Rubin.
Easier Installation and Liquid Cooling
Nvidia reduced the number of cables needed to connect chips to racks. The company touts Vera Rubin as "cable-free compute" and "hot-swappable." Andrew Bell, Nvidia's senior vice president of hardware engineering, discussed the installation improvements. Installation time per rack has been reduced from a couple of hours to a few minutes.
Stay ahead of the AI curve
The most important updates, news, and content — delivered weekly.
No spam. Unsubscribe anytime.
Vera Rubin is 100 percent liquid-cooled. The Vera Rubin NVL72 racks are liquid-cooled platforms. Liquid cooling reduces the energy needed to cool chips compared to air-cooling, a selling point for data center operators facing rising power costs.
The design choices reflect a broader engineering philosophy. Vera abandons chiplet architecture in favor of a single, monolithic chip. Hannah Coutand, Nvidia's CPU product marketing lead, emphasized this approach. Coutand argued chiplets impose "a heavy tax on memory bandwidth and data movement," and that the monolithic design allows data to move more quickly across a single integrated circuit.
This puts Nvidia in contrast with AMD, which is a pioneer of modern chiplet architecture used in x86 processors. x86 processors account for the vast majority of data center CPU revenue. Nvidia builds its data center CPUs on ARM architecture instead.
AMD's Countermove
AMD's annual conference this week is expected to tout next-generation AI and data center chips. On Sunday, AMD revealed details about its Helios AI chip rack. AMD and Nvidia are vying for large-scale, multiyear contracts with hyperscalers and AI labs, including Meta, Amazon, OpenAI, Anthropic, and SpaceXAI.
AMD has significantly grown its share of the market for CPUs used in data centers over the past two years. That growth gives AMD momentum heading into its event, even as Nvidia tries to steal attention with fresh benchmarks.
The timing of Nvidia's promotional push is notable. By releasing performance data days before AMD's showcase, Nvidia is trying to shape the conversation around AI infrastructure before its rival takes the stage.
Competitive Stakes
The competition between AMD and Nvidia extends beyond raw performance. Both companies are chasing the same large customers, and both are pitching complete systems rather than individual chips. Nvidia's move into CPUs signals that it sees the data center as a full-systems opportunity, not just a GPU market.
Nvidia's CPU sales to China could be ready as soon as August, according to reports. That timeline would put the Vera CPU in the market before the full Vera Rubin system ships in the second half of this year.
The company's sensitivity to delays is understandable given the Blackwell episode. Design changes and shipment pushbacks cost time and credibility. With Vera Rubin, Nvidia is emphasizing reliability and ease of deployment through reduced cabling, liquid cooling, and faster installation.
The story was updated on 7/24/26 at 2:23 pm ET to clarify which parts of the Vera Rubin system certain executives discussed.

.jpg)