Qwen2.5-0.5B|1.5B|3B|7B|14B|32B|72B
FreeA Party of Foundation Models!
About Qwen2.5-0.5B|1.5B|3B|7B|14B|32B|72B
Qwen2.5 is a family of open-source, dense, decoder-only language models released by the Qwen Team (Alibaba Cloud) in September 2024. The models are pretrained on up to 18 trillion tokens and are available in sizes of 0.5B, 1.5B, 3B, 7B, 14B, 32B, and 72B parameters. Compared to Qwen2, Qwen2.5 achieves significant improvements in knowledge (MMLU 85+), coding (HumanEval 85+), and mathematics (MATH 80+). It also enhances instruction following, long text generation (up to 8K output tokens), structured data understanding (e.g., tables), and structured JSON output. The models support a context length of up to 128K tokens and are multilingual, covering 29 languages including Chinese, English, French, Spanish, Portuguese, German, Italian, Russian, Japanese, Korean, Vietnamese, Thai, Arabic, and more. All models except the 3B and 72B variants are licensed under Apache 2.0. Additionally, the release includes specialized expert models: Qwen2.5-Coder (trained on 5.5 trillion tokens of code data) and Qwen2.5-Math (supporting Chain-of-Thought, Program-of-Thought, and Tool-Integrated Reasoning in both Chinese and English).
Key Features
Pros & Cons
- State-of-the-art performance among open-source models in knowledge, coding, and math benchmarks
- Wide range of model sizes to suit different computational budgets
- Large context window (128K tokens) for long documents and codebases
- Multilingual support covering 29 languages
- Open-source with permissive Apache 2.0 license for most variants
- Includes specialized expert models for coding and math
- Resilient to diverse system prompts, enhancing role-play and condition-setting
- The 3B and 72B models are not licensed under Apache 2.0, which may limit some use cases
- No native multimodal capabilities (vision, audio) in the Qwen2.5 LLM release (separate Qwen2-VL available)
- Model sizes up to 72B may be too large for local deployment without significant hardware
- Training data specifics and potential biases are not disclosed in detail