AI Models

DeepSeek V4 Model Release: Efficiency Gains

Chinese AI company DeepSeek launched a preview of its V4 model, featuring longer prompt handling and open-source access. The model offers strong performance at low costs, a 1 million token context window, and optimization for Chinese chips like Huawei's Ascend. It builds on R1's success while addressing efficiency and hardware independence.

Neura News

Neura News

Neura Market Editorial

April 24, 20264 min read

Originally reported by technologyreview.com

DeepSeek V4 Model Release: Efficiency Gains

DeepSeek V4 Model Release: Efficiency Gains

Chinese AI company DeepSeek unveiled a preview of its flagship V4 model on Friday. This update allows the system to manage much longer inputs compared to prior versions. The design improves text processing efficiency. As with earlier releases, V4 remains open source, so users can freely download, apply, and adjust it.

This launch represents DeepSeek's biggest step forward since R1 appeared in January 2025. That reasoning model, built with constrained compute, impressed the worldwide AI sector through solid results and low resource needs. It quickly elevated DeepSeek from obscurity to China's leading AI name. R1 also sparked a trend of open-weight models from other Chinese developers.

DeepSeek stayed quiet afterward. This month, however, it introduced "expert" and "flash" options to its web-based model. Those changes fueled talk of a major announcement soon.

The firm stands as a key emblem of China's AI goals. Yet its return to top-tier models follows challenges like staff exits, postponed releases, and attention from US and Chinese regulators.

V4 will not disrupt the field like R1. Still, three factors make it important.

Strong Open-Source Performance at Low Cost

DeepSeek states V4 matches top models while costing far less. Developers and businesses benefit from frontier abilities without high expenses.

Two variants exist on the company's site and app, with API for coders. V4-Pro targets coding and advanced agent work as a bigger system. V4-Flash prioritizes speed and affordability as a compact option. Both include reasoning modes that break down prompts step by step.

Pricing beats rivals. V4-Pro runs at $1.74 per million input tokens and $3.48 per million output tokens, below OpenAI and Anthropic rates. V4-Flash costs $0.14 per million input and $0.28 per million output, among the lowest for elite models. This suits app development well.

Benchmarks show V4-Pro near closed models like Anthropic's Claude-Opus-4.6, OpenAI's GPT-5.4, and Google's Gemini-3.1. Against open rivals such as Alibaba's Qwen-3.5 or Z.ai's GLM-5.1, it leads in coding, math, and STEM tasks. DeepSeek calls it a top open-source release.

V4-Pro excels in agentic coding and multistep challenges per company data. Writing and knowledge tests also rank high.

A technical report cites an internal poll of 85 skilled developers. Over 90% ranked V4-Pro highly for coding.

The model suits frameworks like Claude Code, OpenClaw, and CodeBuddy.

Advances in Context Handling

V4's standout feature is its 1 million token context window. That holds all three Lord of the Rings volumes plus The Hobbit. DeepSeek sets this as standard across services, equaling Gemini and Claude leaders.

The method matters. V4 alters past architectures, mainly attention, which links prompt parts. Long texts raise comparison costs, bottlenecking long-context use.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

V4 selectively attends. It compresses past data, prioritizes relevant sections, and retains recent details fully.

This cuts long-context expenses. For 1 million tokens, V4-Pro needs 27% of V3.2's compute and 10% memory. V4-Flash uses 10% compute and 7% memory.

Such savings aid tools like code assistants scanning full repositories or agents reviewing document sets without memory loss.

DeepSeek researched memory over 18 months via papers on compression and math for extended handling.

Move Toward Chinese Hardware

V4 first optimizes for local chips like Huawei's Ascend. This tests if China can reduce reliance on Nvidia.

Reports noted no early access for Nvidia or AMD, unlike norms. DeepSeek shared previews only with Chinese makers.

Huawei confirmed Ascend 950 supernodes support V4. Users can run customized versions on these chips.

Reuters said officials urged Huawei integration in training. This aligns with self-reliance drives. US controls since 2022 blocked top Nvidia chips, then weaker ones.

China pushes domestic stacks via data center rules, foreign chip limits, quotas, and pairings with Huawei or Cambricon.

Switching challenges Nvidia's software edge. Huawei requires code tweaks and tool rebuilds for stability.

DeepSeek runs V4 inference on Chinese chips. Tsinghua's Liu Zhiyuan notes partial training adaptation; long-context unclear, likely Nvidia-heavy. Sources say Chinese chips lag Nvidia but suit inference over training.

V4-Pro costs may drop with Ascend 950 scale in late 2026.

Success signals parallel AI infrastructure growth.

Related on Neura Market

More from Neura News

AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google has released three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and knowledge work with 17% fewer output tokens and lower costs. The 3.5 Flash-Lite is the fastest in the series at 350 tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber model, available only to governments and trusted partners via CodeMender, focuses on finding and fixing cybersecurity vulnerabilities. Google also noted that Gemini 3.5 Pro is being tested with partners and that pre-training for Gemini 4 has begun.

Jul 21·5 min read
AI Models

Alibaba Qwen-Image-3.0 renders infographics and tiny text in one pass

Alibaba's Qwen team released Qwen-Image-3.0, an image generator designed for practical applications like newspaper layouts and complex infographics. The model processes prompts of up to 4,500 tokens and can render legible text as small as ten pixels, mathematical formulas, and twelve languages in a single pass. It is currently available through invite-only API access, with plans to integrate it into first-party apps like Qwen Chat soon.

Jul 21·4 min read
AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google DeepMind has introduced three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and multimodal performance with 17% fewer output tokens and lower cost. The 3.5 Flash-Lite is the fastest in its series at 350 output tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber, fine-tuned for cybersecurity, will be available exclusively to governments and trusted partners via the CodeMender agent.

Jul 21·6 min read