AI Models

Nvidia Vera chip targets $200B inference market as supply tightens

Nvidia's Q1 revenue beat estimates at $81.62 billion, but CEO Jensen Huang highlighted the Vera chip as a $200 billion market opportunity beyond the $1 trillion from Blackwell and Rubin. Vera revenue is expected to hit $20 billion by fiscal year end, though Huang warned of supply constraints through the Vera Rubin lifecycle. The chip positions Nvidia against hyperscaler custom silicon and Intel/AMD in the inference race.

Neura News

Neura News

Neura Market Editorial

May 21, 20264 min read

Originally reported by artificialintelligence-news.com

Nvidia Vera chip targets $200B inference market as supply tightens

Nvidia reported fiscal first quarter revenue of $81.62 billion on Wednesday, beating analyst estimates of $78.86 billion. The company guided second quarter revenue at $91 billion, well above Wall Street's $86.84 billion forecast. But buried in CEO Jensen Huang's conference call with analysts was a more strategic development than another quarterly beat.

Huang told analysts that Nvidia's new Vera central processors unlock access to a $200 billion market, one that sits entirely outside the $1 trillion the company has already forecast from its Blackwell and Rubin AI GPU lineup between 2025 and 2027. He expects Vera chip revenue to hit $20 billion by the end of this fiscal year. "I expect (Vera) to be the second largest" sales contributor, Huang said during the call.

Why Nvidia needs a second front

The reason Nvidia requires another growth pillar is clear: its biggest customers are building their own chips. Google, Amazon, and Microsoft are collectively expected to pour more than $700 billion into AI infrastructure this year, up sharply from around $400 billion in 2025. They are simultaneously pouring funds into custom silicon to run AI models. Intel and AMD are also touting CPUs as a credible play for inference workloads.

The narrative in the chip industry has shifted from who can train the biggest model to who can serve it cheapest and fastest. Inference is where Nvidia's GPU dominance is most exposed. Training large models remains firmly Nvidia territory, but inference generating answers at scale in real time is increasingly where custom chips from Google's TPU line, Amazon's Trainium and others are making their case.

Vera's target and supply challenges

Nvidia's answer is Vera. The chip, developed in part using technology from Groq, a startup specialising in inference that Nvidia licensed in a deal reportedly worth around $17 billion, targets exactly this workload. The full Vera Rubin platform, which combines the Vera CPU with Rubin GPUs, is set to launch later this year.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

Huang was candid about one problem: supply. "My sense is that we'll be supply-constrained through the entire life of Vera Rubin," he said on the call. It is a telling admission for a product Nvidia is positioning as a major growth pillar. To get ahead of disruptions, Nvidia is spending heavily on the supply chain. The company disclosed that its supply commitments rose to $119 billion in Q1, up from $95.2 billion the previous quarter, a significant jump that reflects both confidence in demand and anxiety about a global memory chip crunch.

Financial moves and market reaction

Nvidia also announced an $80 billion share repurchase programme and raised its quarterly cash dividend to 25 cents per share from 1 cent. These moves signal financial confidence even as Huang warned of tightening supply.

Despite the beats, Nvidia shares fell 1.6% in extended trading after the results. eMarketer analyst Jacob Bourne captured the mood: "Nvidia delivered another beat, but at this point that's essentially priced in as it keeps beating quarter after quarter. The lingering question is whether it can convince investors the AI buildout has durability into 2027 and 2028, especially as the narrative shifts toward inference workloads and competing silicon from Google, Amazon, AMD, and Intel."

Huang pushed back with numbers of his own. He pointed to a growing sub-segment of AI-specific cloud customers whose spend is now roughly equal to the hyperscalers, but growing faster quarter over quarter. "We should be growing faster than hyperscale capex," he said.

The Vera chip is central to that argument. Whether the supply chain cooperates is a different question entirely.

, -

Related on Neura Market:

More from Neura News

AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google has released three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and knowledge work with 17% fewer output tokens and lower costs. The 3.5 Flash-Lite is the fastest in the series at 350 tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber model, available only to governments and trusted partners via CodeMender, focuses on finding and fixing cybersecurity vulnerabilities. Google also noted that Gemini 3.5 Pro is being tested with partners and that pre-training for Gemini 4 has begun.

Jul 21·5 min read
AI Models

Alibaba Qwen-Image-3.0 renders infographics and tiny text in one pass

Alibaba's Qwen team released Qwen-Image-3.0, an image generator designed for practical applications like newspaper layouts and complex infographics. The model processes prompts of up to 4,500 tokens and can render legible text as small as ten pixels, mathematical formulas, and twelve languages in a single pass. It is currently available through invite-only API access, with plans to integrate it into first-party apps like Qwen Chat soon.

Jul 21·4 min read
AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google DeepMind has introduced three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and multimodal performance with 17% fewer output tokens and lower cost. The 3.5 Flash-Lite is the fastest in its series at 350 output tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber, fine-tuned for cybersecurity, will be available exclusively to governments and trusted partners via the CodeMender agent.

Jul 21·6 min read