DeepSeek-V3 Technical Report
FreeStrong MoE language model with 671B parameters and 37B activated per token
About DeepSeek-V3 Technical Report
DeepSeek-V3 is a powerful Mixture-of-Experts (MoE) language model developed by DeepSeek-AI, featuring 671 billion total parameters with 37 billion activated per token. It builds upon the architectures validated in DeepSeek-V2, employing Multi-head Latent Attention (MLA) and DeepSeekMoE to achieve efficient inference and cost-effective training. The model introduces an auxiliary-loss-free strategy for load balancing and a multi-token prediction training objective to enhance performance. Pre-trained on 14.8 trillion tokens of diverse, high-quality data, it undergoes supervised fine-tuning and reinforcement learning to fully harness its capabilities. Comprehensive evaluations show that DeepSeek-V3 outperforms other open-source models and rivals leading closed-source models across a wide range of benchmarks.
Key Features
Pros & Cons
- State-of-the-art performance among open-source models, matching closed-source competitors
- Efficient inference via MoE with only 37B activated parameters
- Cost-effective training due to architectural innovations
- Novel load balancing strategy without auxiliary losses
- Multi-token prediction objective enhances learning effectiveness