Solving math word problems with processand outcome-based feedback
FreeComparing process- and outcome-based feedback for math reasoning in language models
About Solving math word problems with processand outcome-based feedback
This paper presents a comprehensive comparison between process-based and outcome-based supervision approaches for training language models to solve math word problems (GSM8K). The study finds that pure outcome-based supervision achieves similar final-answer error rates with less label supervision, but correct reasoning steps require process-based supervision or learned reward models that emulate process-based feedback. The authors improve previous state-of-the-art results, reducing final-answer error from 16.8% to 12.7% and reasoning error among final-answer-correct solutions from 14.0% to 3.4%. The work highlights the trade-offs between supervision types for reasoning tasks and has implications for education and real-world applications where reasoning accuracy matters.
Key Features
Pros & Cons
- Outcome-based supervision requires fewer labels while maintaining similar final-answer accuracy
- Process-based supervision significantly reduces reasoning errors
- Comprehensive comparison on a standard benchmark (GSM8K)
- Improves upon previous best results on GSM8K
- Process-based supervision requires more detailed labels (reasoning steps)
- Learned reward models may not fully capture process-based feedback quality
- Results are specific to math word problems; generalizability to other reasoning tasks is not explored