Qwen2-Math-1.5B|7B|72B logo

Qwen2-Math-1.5B|7B|72B

Free

Math-specific LLMs outperforming GPT-4o and other SOTA models

FreeFree tier
Inputs: textOutputs: text
Type
Open Source
Company
Alibaba Cloud

About Qwen2-Math-1.5B|7B|72B

Qwen2-Math is a specialized series of large language models designed for mathematical reasoning, built upon the Qwen2 architecture. It includes base and instruction-tuned variants in 1.5B, 7B, and 72B parameter sizes. The models are pretrained on a meticulously designed mathematics-specific corpus containing high-quality web texts, books, code, exam questions, and synthesized data. Instruction tuning uses a math-specific reward model and Group Relative Policy Optimization (GRPO). Qwen2-Math-Instruct achieves state-of-the-art performance on a wide range of English and Chinese math benchmarks, outperforming GPT-4o, Claude-3.5-Sonnet, Gemini-1.5-Pro, and Llama-3.1-405B on many tasks, including GSM8K, MATH, OlympiadBench, AIME 2024, and AMC 2023. The model is open-source and freely available on GitHub, Hugging Face, and ModelScope.

Key Features

Available in 1.5B, 7B, and 72B parameter sizes
Base model pretrained on large-scale high-quality mathematics corpus (web texts, books, code, exam questions, synthesized data)
Instruction-tuned variant (Instruct) trained with reward model and Group Relative Policy Optimization (GRPO)
Outperforms GPT-4o, Claude-3.5-Sonnet, Gemini-1.5-Pro, and Llama-3.1-405B on math benchmarks
Strong performance on competition-level problems (AIME 2024, AMC 2023) and diverse math exams (OlympiadBench, CollegeMath, GaoKao)
Supports few-shot chain-of-thought prompting and zero-shot evaluation
Open-source with models available on GitHub, Hugging Face, and ModelScope

Pros & Cons

Pros
  • Achieves state-of-the-art results on multiple math benchmarks, surpassing leading closed-source models
  • Specialized training on mathematics corpus enhances reasoning capabilities
  • Open-source and freely available for research and commercial use
  • Multiple model sizes allow deployment on varying hardware constraints
  • Comprehensive evaluation on both English and Chinese math datasets
Cons
  • Currently only supports English; bilingual (English and Chinese) version is forthcoming
  • Largest 72B model requires substantial computational resources
  • Limited to mathematical tasks; not designed for general-purpose use
  • Performance may vary on non-math reasoning or creative tasks

Best For

Solving arithmetic and mathematical problems across various difficulty levelsMath competition preparation and problem solving (IMO, AIME, AMC)Educational tutoring for mathematics (high school, college, olympiad)Automated math reasoning and proof generationEvaluation on standardized math tests (GSM8K, MATH, MMLU-STEM, CMATH, GaoKao)

FAQ

What is Qwen2-Math?
Qwen2-Math is a series of math-specific large language models built on the Qwen2 architecture, designed to excel at arithmetic and mathematical problem solving. It includes base and instruction-tuned versions in 1.5B, 7B, and 72B sizes.
How does Qwen2-Math compare to GPT-4o?
According to the developers, Qwen2-Math-72B-Instruct outperforms GPT-4o, as well as Claude-3.5-Sonnet, Gemini-1.5-Pro, and Llama-3.1-405B, on a series of math benchmarks.
Is Qwen2-Math open-source?
Yes, Qwen2-Math is open-source and available on GitHub, Hugging Face, and ModelScope under the Qwen organization.
Does Qwen2-Math support languages other than English?
The current release mainly supports English. The team plans to release bilingual (English and Chinese) math models soon.
What sizes are available?
Qwen2-Math is available in three sizes: 1.5B, 7B, and 72B parameters.