Qwen2-Math-1.5B|7B|72B
FreeMath-specific LLMs outperforming GPT-4o and other SOTA models
About Qwen2-Math-1.5B|7B|72B
Qwen2-Math is a specialized series of large language models designed for mathematical reasoning, built upon the Qwen2 architecture. It includes base and instruction-tuned variants in 1.5B, 7B, and 72B parameter sizes. The models are pretrained on a meticulously designed mathematics-specific corpus containing high-quality web texts, books, code, exam questions, and synthesized data. Instruction tuning uses a math-specific reward model and Group Relative Policy Optimization (GRPO). Qwen2-Math-Instruct achieves state-of-the-art performance on a wide range of English and Chinese math benchmarks, outperforming GPT-4o, Claude-3.5-Sonnet, Gemini-1.5-Pro, and Llama-3.1-405B on many tasks, including GSM8K, MATH, OlympiadBench, AIME 2024, and AMC 2023. The model is open-source and freely available on GitHub, Hugging Face, and ModelScope.
Key Features
Pros & Cons
- Achieves state-of-the-art results on multiple math benchmarks, surpassing leading closed-source models
- Specialized training on mathematics corpus enhances reasoning capabilities
- Open-source and freely available for research and commercial use
- Multiple model sizes allow deployment on varying hardware constraints
- Comprehensive evaluation on both English and Chinese math datasets
- Currently only supports English; bilingual (English and Chinese) version is forthcoming
- Largest 72B model requires substantial computational resources
- Limited to mathematical tasks; not designed for general-purpose use
- Performance may vary on non-math reasoning or creative tasks