Industry

GPU financiers shift to inference chips in $400M deal

General Compute, an AI inference cloud startup, secured a $400 million loan from Upper90, marking what may be the first deal to use inference-specific chips as collateral. The financing signals growing market demand for cost-efficient AI infrastructure that runs open-source models, as investors seek alternatives to expensive GPU-based systems.

Neura News

Neura News

Neura Market Editorial

July 17, 20264 min read

Originally reported by techcrunch.com

GPU financiers shift to inference chips in $400M deal

General Compute, a startup focused on AI inference cloud services, has obtained a $400 million loan from Upper90, a technology investment firm. This transaction could represent the first instance where inference-specific chips were used as collateral. These chips are designed to run already trained AI models quickly and efficiently, unlike the more costly chips needed to build the models initially.

Market response to AI pricing concerns

The financing reflects how markets are reacting to worries about the price of AI tools and tokens. Investors are turning to infrastructure that can run open-source models at a lower cost compared to the newest large language models from frontier labs.

General Compute was founded by CEO Finn Puklowski. The company raised a $15 million seed round in May to build an inference neocloud using silicon from SambaNova, a chipmaker backed by Intel. Neoclouds are purpose built for AI workloads, unlike the general purpose infrastructure offered by traditional hyperscalers such as AWS or Azure.

Chip design advantages

The company's SN50 chips are specifically designed for inference. They are power efficient and do not require expensive water cooling systems. This allows them to be deployed more quickly than GPUs across a wider variety of data centers. General Compute claims the new chips will deliver 16 times faster inference than GPU based clouds.

A key challenge for a brand new company is acquiring a large number of these chips.

Financing strategy

Upper90 co-founder and CEO Billy Libby, a former Goldman Sachs quantitative trader, had a playbook for this situation. In 2021, his firm financed GPU purchases by Crusoe, the energy focused data center startup. Libby believes that was the first loan against the value of advanced chips.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

Traditional lenders avoided such deals at the time due to risks and uncertainties around GPU depreciation. However, as CoreWeave turned chips backed loans into a business model and later the foundation of a blockbuster IPO, this type of financing has become common.

"When we financed Nvidia GPUs as the first group to do that, the market was inefficient," Libby told TechCrunch. "We could really put together something as an early participant, and kind of get compensated for the risk."

Now that GPUs are comparatively well understood and possibly overbought, Upper90 is turning to companies like General Compute to ride the next wave of the AI boom. "We think open source models are going to be important, and we went and looked for a player last year that was in inference," Libby said. "Everyone doesn't need a supercomputer, but they do need inference and AI."

Growing thesis

That thesis has been gaining strength. Companies that provide access to open models, such as OpenRouter and Fireworks, have raised new rounds at huge valuations. New models like Kimi's K3, released just this week, have proven to compete with the latest releases from Anthropic and OpenAI on coding benchmarks. New chipmakers like Groq and Cerebras have drawn interest from acquirers and public markets alike.

General Compute's ability to access chips outside of Nvidia's ecosystem matters for the same reason. TensorWave, another AI infrastructure company, is making a similar bet on a partnership with AMD. As more alternatives to Nvidia emerge, compute providers that are not locked into Nvidia deals may have an advantage in providing cost efficient inference.

"There are a bunch of chips that are starting to scale that have amazing total cost of ownership, or that can operate much faster than Nvidia, but there's not too many buyers for them," Puklowski said. "By getting together with Upper90, this is not just, 'a cool startup got some money to buy some compute.' Like, this is the first signal of capital organizing itself and the fragmenting of Nvidia's monopolistic dominance."

Related on Neura Market

More from Neura News

AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google has released three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and knowledge work with 17% fewer output tokens and lower costs. The 3.5 Flash-Lite is the fastest in the series at 350 tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber model, available only to governments and trusted partners via CodeMender, focuses on finding and fixing cybersecurity vulnerabilities. Google also noted that Gemini 3.5 Pro is being tested with partners and that pre-training for Gemini 4 has begun.

Jul 21·5 min read
AI Models

Alibaba Qwen-Image-3.0 renders infographics and tiny text in one pass

Alibaba's Qwen team released Qwen-Image-3.0, an image generator designed for practical applications like newspaper layouts and complex infographics. The model processes prompts of up to 4,500 tokens and can render legible text as small as ten pixels, mathematical formulas, and twelve languages in a single pass. It is currently available through invite-only API access, with plans to integrate it into first-party apps like Qwen Chat soon.

Jul 21·4 min read
AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google DeepMind has introduced three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and multimodal performance with 17% fewer output tokens and lower cost. The 3.5 Flash-Lite is the fastest in its series at 350 output tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber, fine-tuned for cybersecurity, will be available exclusively to governments and trusted partners via the CodeMender agent.

Jul 21·6 min read