AI Models

DeepSeek V4 Models Close Gap to Frontier AI

DeepSeek released preview versions of DeepSeek V4 Flash and V4 Pro, mixture-of-experts models with 1 million token contexts. The Pro version features 1.6 trillion parameters, the largest open-weight model. They match top models on reasoning and coding benchmarks while offering lower prices.

Neura News

Neura News

Neura Market Editorial

April 24, 20263 min read

Originally reported by techcrunch.com

DeepSeek V4 Models Close Gap to Frontier AI

DeepSeek V4 Models Close Gap to Frontier AI

Chinese AI laboratory DeepSeek unveiled preview editions of its latest large language model, DeepSeek V4. This includes two variants: V4 Flash and V4 Pro. The release updates last year's V3.2 model and the R1 reasoning model, both of which gained widespread attention in the AI community.

Model Architecture and Scale

DeepSeek describes V4 Flash and V4 Pro as mixture-of-experts systems. Each supports a context window of 1 million tokens. That capacity handles extensive codebases or lengthy documents in single prompts. The mixture-of-experts design activates a subset of parameters for each task. This reduces costs during inference.

The V4 Pro holds 1.6 trillion total parameters, with 49 billion active. It stands as the largest open-weight model released so far. It surpasses Moonshot AI's Kimi K 2.6 at 1.1 trillion parameters, MiniMax's M1 with 456 billion, and exceeds DeepSeek V3.2's 671 billion by more than double. V4 Flash, the compact option, contains 284 billion parameters, 13 billion active.

DeepSeek, founded as a key player in China's AI efforts, focuses on open-source releases to compete globally. Past models like V3.2 demonstrated strong performance at low costs, drawing developers and researchers.

Performance on Benchmarks

According to DeepSeek, the new models outperform V3.2 in efficiency and results thanks to design upgrades. They have nearly matched leading open and closed models on reasoning tests. The V4-Pro-Max variant beats other open-source competitors across reasoning evaluations. It also tops OpenAI's GPT-5.2 and Gemini 3.0 Pro in certain tasks.

On coding competition benchmarks, both V4 models deliver results similar to GPT-5.4. However, they trail slightly on knowledge assessments against GPT-5.4 and Google's Gemini 3.1 Pro. DeepSeek notes this indicates a development path about 3 to 6 months behind top frontier models.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

Unlike many proprietary rivals, V4 Flash and V4 Pro handle text exclusively. They lack features for audio, video, or image processing found in some closed systems.

Pricing Advantages

DeepSeek V4 offers costs far below current frontier options. V4 Flash charges $0.14 per million input tokens and $0.28 per million output tokens. Those rates beat GPT-5.4 Nano, Gemini 3.1 Flash, GPT-5.4 Mini, and Claude Haiku 4.5.

V4 Pro prices at $0.145 per million input tokens and $3.48 per million output tokens. It undercuts Gemini 3.1 Pro, GPT-5.5, Claude Opus 4.7, and GPT-5.4. Such affordability stems from the efficient architecture and open-weight approach, making high capability accessible.

Broader Context

The announcement arrived one day after U.S. authorities charged China with large-scale theft of American AI intellectual property via thousands of proxy accounts. Separately, Anthropic and OpenAI have accused DeepSeek of distilling, a process that copies elements from their models.

DeepSeek continues to push boundaries in open AI development from China. Its models provide alternatives to dominant U.S. providers, emphasizing scale, cost, and benchmark parity.

Related on Neura Market

More from Neura News

AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google has released three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and knowledge work with 17% fewer output tokens and lower costs. The 3.5 Flash-Lite is the fastest in the series at 350 tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber model, available only to governments and trusted partners via CodeMender, focuses on finding and fixing cybersecurity vulnerabilities. Google also noted that Gemini 3.5 Pro is being tested with partners and that pre-training for Gemini 4 has begun.

Jul 21·5 min read
AI Models

Alibaba Qwen-Image-3.0 renders infographics and tiny text in one pass

Alibaba's Qwen team released Qwen-Image-3.0, an image generator designed for practical applications like newspaper layouts and complex infographics. The model processes prompts of up to 4,500 tokens and can render legible text as small as ten pixels, mathematical formulas, and twelve languages in a single pass. It is currently available through invite-only API access, with plans to integrate it into first-party apps like Qwen Chat soon.

Jul 21·4 min read
AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google DeepMind has introduced three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and multimodal performance with 17% fewer output tokens and lower cost. The 3.5 Flash-Lite is the fastest in its series at 350 output tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber, fine-tuned for cybersecurity, will be available exclusively to governments and trusted partners via the CodeMender agent.

Jul 21·6 min read