AI Models

Deepseek makes 75% discount permanent, undercuts GPT-5.5 by 34x on output pricing

Deepseek has made its 75 percent discount on the flagship Deepseek V4 Pro model permanent, slashing output token prices to $0.87 per million versus $30 for GPT-5.5. The move intensifies the price war between Chinese and Western AI labs, with Deepseek offering output tokens at roughly 34.5 times less than OpenAI's top model. The permanent pricing, originally set to expire May 31, 2026, also undercuts Anthropic's Opus 4.7 by a similar margin on output tokens.

Neura News

Neura News

Neura Market Editorial

May 23, 20263 min read

Originally reported by the-decoder.com

Deepseek makes 75% discount permanent, undercuts GPT-5.5 by 34x on output pricing

Deepseek has turned its 75 percent discount on the flagship Deepseek V4 Pro model into a permanent price cut, the company announced on X. The promotion was originally scheduled to expire on May 31, 2026.

Under the permanent pricing, one million input tokens without cache cost $0.435, while one million output tokens cost $0.87. Cache hits push the input price even lower. In comparison, GPT 5.5 charges $5 per million input tokens and $30 per million output tokens, while Opus 4.7 sits at $5 for input and $25 for output.

Pricing comparison at a glance

The table below shows how Deepseek's models stack up against competitors on a per-token basis.

| Model | Input per 1M tokens | Input cache hit | Output per 1M tokens | |, , , -|, , , , , , , , , , -|, , , , , , , , |, , , , , , , , , , , | | Deepseek-V4-Pro | $0.435 | $0.003625 | $0.87 | | Deepseek-V4-Flash | $0.14 | $0.0028 | $0.28 | | GPT-5.5 | $5.00 | $0.50 | $30.00 | | GPT-5.5 (Long Context, >272K) | $10.00 | $1.00 | $45.00 | | Opus 4.7 | $5.00 | $0.50 | $25.00 |

That makes Deepseek's flagship about 11.5 times cheaper than GPT 5.5 on standard input pricing. The gap is much wider on output, where Deepseek V4 Pro is about 34.5 times cheaper. Against GPT 5.5 long context pricing above 272K tokens, Deepseek V4 Pro is about 23 times cheaper on input and about 51.7 times cheaper on output. Deepseek V4 Flash is cheaper still.

Both Deepseek models offer a one million token context window and up to 384,000 output tokens. Deepseek also supports both OpenAI and Anthropic API formats, making it easier for developers to switch.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

Token prices only tell half the story

Raw per-token pricing is only part of the picture, though. Token consumption per task matters just as much. Think of it like gas prices: a low price per gallon does not help if your engine guzzles fuel.

A good example is Google's Gemini Flash 3.5. On paper, it is cheaper and performs similarly to the previous Pro model 3.1, but it burns through far more tokens, making it potentially pricier in practice. Anthropic's Opus 4.7 looks cheaper on paper than GPT-5.5 too, but uses more tokens than its predecessor. GPT-5.5, on the other hand, consumes fewer tokens than GPT-5.4. Still, both models ended up 30 to 90 percent more expensive than the models they replaced.

Deepseek V4 clearly trails the top frontier models GPT-5.5 and Opus 4.7 in raw performance. How much exactly depends on the task, and benchmarks only tell half the story. Only real-world use will tell. But the price gap is massive, especially for agentic AI systems that chew through many times more tokens than a standard chatbot.

As AI usage grows, companies are getting more price-sensitive. As long as ROI on AI spending remains hard to measure, many firms may shift strategy: away from the best model and toward the cheapest one that is still good enough.

Deepseek is entering its first funding round, but it faces nowhere near the revenue pressure that OpenAI and Anthropic do. Both of those labs are also heading toward IPOs.

Related on Neura Market

More from Neura News

AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google has released three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and knowledge work with 17% fewer output tokens and lower costs. The 3.5 Flash-Lite is the fastest in the series at 350 tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber model, available only to governments and trusted partners via CodeMender, focuses on finding and fixing cybersecurity vulnerabilities. Google also noted that Gemini 3.5 Pro is being tested with partners and that pre-training for Gemini 4 has begun.

Jul 21·5 min read
AI Models

Alibaba Qwen-Image-3.0 renders infographics and tiny text in one pass

Alibaba's Qwen team released Qwen-Image-3.0, an image generator designed for practical applications like newspaper layouts and complex infographics. The model processes prompts of up to 4,500 tokens and can render legible text as small as ten pixels, mathematical formulas, and twelve languages in a single pass. It is currently available through invite-only API access, with plans to integrate it into first-party apps like Qwen Chat soon.

Jul 21·4 min read
AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google DeepMind has introduced three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and multimodal performance with 17% fewer output tokens and lower cost. The 3.5 Flash-Lite is the fastest in its series at 350 output tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber, fine-tuned for cybersecurity, will be available exclusively to governments and trusted partners via the CodeMender agent.

Jul 21·6 min read