Qwen2.5-0.5B|1.5B|3B|7B|14B|32B|72B logo

Qwen2.5-0.5B|1.5B|3B|7B|14B|32B|72B

Free

A Party of Foundation Models!

FreeFree tier
Inputs: textOutputs: text, code
Type
Open Source
Company
Alibaba Cloud (Qwen Team)

About Qwen2.5-0.5B|1.5B|3B|7B|14B|32B|72B

Qwen2.5 is a family of open-source, dense, decoder-only language models released by the Qwen Team (Alibaba Cloud) in September 2024. The models are pretrained on up to 18 trillion tokens and are available in sizes of 0.5B, 1.5B, 3B, 7B, 14B, 32B, and 72B parameters. Compared to Qwen2, Qwen2.5 achieves significant improvements in knowledge (MMLU 85+), coding (HumanEval 85+), and mathematics (MATH 80+). It also enhances instruction following, long text generation (up to 8K output tokens), structured data understanding (e.g., tables), and structured JSON output. The models support a context length of up to 128K tokens and are multilingual, covering 29 languages including Chinese, English, French, Spanish, Portuguese, German, Italian, Russian, Japanese, Korean, Vietnamese, Thai, Arabic, and more. All models except the 3B and 72B variants are licensed under Apache 2.0. Additionally, the release includes specialized expert models: Qwen2.5-Coder (trained on 5.5 trillion tokens of code data) and Qwen2.5-Math (supporting Chain-of-Thought, Program-of-Thought, and Tool-Integrated Reasoning in both Chinese and English).

Key Features

Pretrained on up to 18 trillion tokens for enhanced knowledge and capabilities
Available in 7 sizes: 0.5B, 1.5B, 3B, 7B, 14B, 32B, and 72B parameters
Achieves MMLU 85+, HumanEval 85+, and MATH 80+ benchmarks
Improved instruction following, long text generation (8K tokens), and structured output (JSON)
Supports up to 128K input tokens and generates up to 8K tokens
Multilingual support for 29 languages including English, Chinese, French, Spanish, German, Arabic, Japanese, Korean, and more
Includes specialized expert models: Qwen2.5-Coder (code-focused) and Qwen2.5-Math (mathematics-focused)
Qwen2.5-Coder trained on 5.5 trillion tokens of code-related data
Qwen2.5-Math supports Chain-of-Thought, Program-of-Thought, and Tool-Integrated Reasoning
Mostly licensed under Apache 2.0 (except 3B and 72B variants)

Pros & Cons

Pros
  • State-of-the-art performance among open-source models in knowledge, coding, and math benchmarks
  • Wide range of model sizes to suit different computational budgets
  • Large context window (128K tokens) for long documents and codebases
  • Multilingual support covering 29 languages
  • Open-source with permissive Apache 2.0 license for most variants
  • Includes specialized expert models for coding and math
  • Resilient to diverse system prompts, enhancing role-play and condition-setting
Cons
  • The 3B and 72B models are not licensed under Apache 2.0, which may limit some use cases
  • No native multimodal capabilities (vision, audio) in the Qwen2.5 LLM release (separate Qwen2-VL available)
  • Model sizes up to 72B may be too large for local deployment without significant hardware
  • Training data specifics and potential biases are not disclosed in detail

Best For

General conversational AI and chatbots with improved role-play and condition-settingCode generation, understanding, and completion tasksMathematical reasoning and problem-solving with multiple reasoning strategiesLong-form content generation (up to 8K tokens)Structured data extraction and analysis (tables, JSON)Multilingual applications across 29 languagesInstruction-following and task automation with structured outputs

FAQ

What model sizes are available for Qwen2.5?
Qwen2.5 is available in 7 sizes: 0.5B, 1.5B, 3B, 7B, 14B, 32B, and 72B parameters. All except 3B and 72B are licensed under Apache 2.0.
What is the context window size for Qwen2.5?
Qwen2.5 supports up to 128K input tokens and can generate up to 8K output tokens.
Which languages does Qwen2.5 support?
Qwen2.5 supports multilingual applications across over 29 languages, including English, Chinese, French, Spanish, Portuguese, German, Italian, Russian, Japanese, Korean, Vietnamese, Thai, Arabic, and more.
Are there specialized models for coding and math?
Yes, the Qwen2.5 release includes Qwen2.5-Coder (trained on 5.5 trillion tokens of code-related data) and Qwen2.5-Math (supports Chain-of-Thought, Program-of-Thought, and Tool-Integrated Reasoning).
How does Qwen2.5 compare to Qwen2?
Qwen2.5 achieves significantly more knowledge (MMLU 85+), greatly improved coding (HumanEval 85+) and math (MATH 80+), better instruction following, longer text generation, improved structured data understanding, and enhanced JSON output.
What are the licensing terms for Qwen2.5 models?
All models except the 3B and 72B variants are licensed under Apache 2.0. The license files are available in the respective Hugging Face repositories.