Industry

Nvidia Aims to Dominate Every Chip in AI Data Centers

Nvidia is promoting its new Vera Rubin chip system, which combines CPUs and GPUs into a single platform. The company is positioning itself as a supplier of complete AI systems, not just GPUs. New benchmarks show significant performance and efficiency gains, with early customer OpenAI already using a rack.

Neura News

Neura News

Neura Market Editorial

July 21, 20265 min read

Originally reported by wired.com

Nvidia Aims to Dominate Every Chip in AI Data Centers

Nvidia is promoting its new Vera Rubin chip system this week, sharing fresh performance benchmarks for the GPU and CPU combination ahead of rival AMD's annual product event in San Francisco on Thursday.

During a detailed technical workshop last week at the company's headquarters in Santa Clara, California, Nvidia executives spoke to a small group of journalists about the chip system's improved power and efficiency. The main message: Nvidia, long known for making GPUs, is increasingly trying to become a supplier of CPUs that can power AI agents.

Vera Rubin: The Next Generation

Vera Rubin is Nvidia's successor to its hybrid superchip system Grace Blackwell. It represents the centerpiece of the company's near-term future powering the AI industry. The system is designed to offer one CPU for every two GPUs. In a single Vera Rubin NVL 72 super chip system, there are 36 Vera CPUs for every 72 Rubin GPUs. Nvidia is also selling the Vera CPU as a standalone product, and it has reportedly told Chinese customers these could be ready as soon as August.

Nvidia executives emphasized that the new Vera Rubin NVL72 racks, which are stacks of chips packed into a single liquid-cooled platform, are much more "plug-and-play" than some earlier products. During a brief tour of a Nvidia data center lab in Silicon Valley, executives shared that OpenAI already has one Vera Rubin rack in use.

Nvidia CEO Jensen Huang did not appear at the workshop in Santa Clara last week. He was in Japan announcing the chipmaker's new partnerships with several Japanese firms to develop AI for robotics. The briefings were instead led by Ian Buck, Nvidia's longtime vice president of accelerated computing and the architect behind the company's CUDA software.

"We're on a roadmap to crank out new architectures, not just GPUs but CPUs," Buck told reporters. "We're going to keep innovating, because it's do this or die in Silicon Valley."

The meetings were held in Huang's executive briefing center. Multiple desks nearby were piled with bags of Taiwanese snacks that the CEO brought back from his recent trip to Computex, a massive annual semiconductor trade show in Taipei, an Nvidia spokesperson told WIRED.

Performance and Efficiency Gains

Nvidia claims that the Vera Rubin NVL72 system will process ten times as many tokens per watt as the company's Grace Blackwell super chip. The company says its Vera CPU is also faster at processing agentic AI tasks compared to rival CPUs from AMD and Intel, though the tests it ran to support those benchmarks appear to have used slightly older generations of its competitors' CPUs. Localized memory subsystems on the new chips will also offer nearly three times as much memory bandwidth as Blackwell, which will likely be an appealing feature to many companies amid an ongoing shortage of high bandwidth memory.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

Nvidia says it has also significantly reduced the number of cables needed to connect its chips to racks in multi-rack server systems. The company is touting Vera Rubin as "cable-free compute" and "hot-swappable." This means customers can theoretically reduce the time it takes to install each rack from a couple of hours to a few minutes, a point raised by both Buck and Andrew Bell, Nvidia's senior vice president of hardware engineering. The new chip system is 100 percent liquid-cooled, which can reduce the amount of energy needed to cool the chips, since air-cooling is more energy intensive.

Release Timeline and Competition

Ever since Nvidia unveiled Vera Rubin in the spring of 2025, the company has been slowly releasing more details about the chip system while insisting it will be released on schedule. Huang has repeatedly said Vera Rubin is ramping to "full production" and will ship in the second half of this year, with early customers including Microsoft, OpenAI, and Oracle.

Nvidia is particularly sensitive to any suggestion of delays after its previous-generation Blackwell chips reportedly overheated when connected together in the company's customized server racks, forcing it to make design changes and push back shipments.

Nvidia's marketing push for Vera Rubin is happening just ahead of rival AMD's annual conference, where executives are expected to tout its next-generation AI and data center chips. On Sunday, AMD revealed more details about its Helios AI chip rack, which is designed to compete with Nvidia's new wares. Both AMD and Nvidia have been vying for large-scale, multi-year contracts to supply AI hyperscalers like Meta and Amazon and AI labs like OpenAI, Anthropic, and SpaceXAI with chips.

Over the past two years, AMD has significantly grown its share of the market for CPUs used in data centers. The company has long been recognized as a pioneer of the modern chiplet architecture used in x86 processors, which still account for the vast majority of data center CPU revenue. Nvidia, by contrast, builds its data center CPUs on ARM, an alternative chip architecture known for its power efficiency.

A Monolithic Design

Nvidia executives Buck and Hannah Coutand, who runs product marketing for Nvidia DGX Cloud, both emphasized that Vera Rubin abandons the chiplet architecture used by many modern processors in favor of a single, monolithic chip. Coutand argued that stitching together multiple chiplets imposes "a heavy tax on memory bandwidth and data movement," whereas the monolithic design of Vera Rubin allows data to move more quickly across a single integrated circuit.

Related on Neura Market

More from Neura News

AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google has released three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and knowledge work with 17% fewer output tokens and lower costs. The 3.5 Flash-Lite is the fastest in the series at 350 tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber model, available only to governments and trusted partners via CodeMender, focuses on finding and fixing cybersecurity vulnerabilities. Google also noted that Gemini 3.5 Pro is being tested with partners and that pre-training for Gemini 4 has begun.

Jul 21·5 min read
AI Models

Alibaba Qwen-Image-3.0 renders infographics and tiny text in one pass

Alibaba's Qwen team released Qwen-Image-3.0, an image generator designed for practical applications like newspaper layouts and complex infographics. The model processes prompts of up to 4,500 tokens and can render legible text as small as ten pixels, mathematical formulas, and twelve languages in a single pass. It is currently available through invite-only API access, with plans to integrate it into first-party apps like Qwen Chat soon.

Jul 21·4 min read
AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google DeepMind has introduced three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and multimodal performance with 17% fewer output tokens and lower cost. The 3.5 Flash-Lite is the fastest in its series at 350 output tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber, fine-tuned for cybersecurity, will be available exclusively to governments and trusted partners via the CodeMender agent.

Jul 21·6 min read