Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement logo

Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Free

Self-improving math expert model series with advanced reasoning

FreeFree tier
Inputs: textOutputs: text
Type
Open Source

About Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Qwen2.5-Math is a series of math-specialized large language models (1.5B, 7B, and 72B parameters) developed by the Qwen team (Alibaba Group). The core innovation is a self-improvement pipeline spanning pre-training, post-training, and inference. During pre-training, Qwen2-Math-Instruct generates large-scale high-quality mathematical data. In post-training, a reward model (RM) is trained via massive sampling from Qwen2-Math-Instruct and used for iterative evolution of supervised fine-tuning (SFT) data; the final SFT model undergoes reinforcement learning with the ultimate RM. At inference, the RM guides sampling. Qwen2.5-Math-Instruct supports both Chinese and English and possesses advanced mathematical reasoning capabilities including Chain-of-Thought (CoT) and Tool-Integrated Reasoning (TIR). The models are evaluated on 10 mathematics datasets in English and Chinese, covering grade school to competition problems (e.g., GSM8K, MATH, GaoKao, AMC23, AIME24). The technical report details the methodology and results.

Key Features

Self-improvement pipeline from pre-training to inference
Large-scale mathematical data generation using Qwen2-Math-Instruct
Reward model trained via massive sampling for iterative SFT data evolution
Reinforcement learning on final SFT model
Supports Chain-of-Thought (CoT) and Tool-Integrated Reasoning (TIR)
Bilingual support: Chinese and English
Available in 1.5B, 7B, and 72B parameter sizes
Evaluated on 10 mathematics datasets including GSM8K, MATH, GaoKao, AMC23, AIME24

Pros & Cons

Pros
  • State-of-the-art performance on diverse math benchmarks
  • Innovative self-improvement training methodology
  • Open-source model available for community use
  • Supports both chain-of-thought and tool-integrated reasoning
Cons
  • Specialized for mathematics; limited general language capabilities
  • Larger models (72B) require significant computational resources
  • Only available in Chinese and English

Best For

Solving grade-school to competition-level math problemsEducational tools for mathematics tutoringAutomated math reasoning for research and applicationsBilingual math problem solving (Chinese and English)

FAQ

What is Qwen2.5-Math?
Qwen2.5-Math is a series of math-specialized large language models developed by the Qwen team. It includes models of sizes 1.5B, 7B, and 72B parameters, focusing on mathematical reasoning via a self-improvement pipeline.
What reasoning methods does Qwen2.5-Math-Instruct support?
It supports Chain-of-Thought (CoT) reasoning and Tool-Integrated Reasoning (TIR), enabling step-by-step problem solving and use of external tools.
On which datasets was Qwen2.5-Math evaluated?
The models were evaluated on 10 mathematics datasets including GSM8K, MATH, GaoKao, AMC23, and AIME24, covering grade school to competition levels in both English and Chinese.
Is Qwen2.5-Math open-source?
Yes, the technical report indicates it is an open-source model series. The models are available for use.