Preprint
Machine Learning

Towards Robust Mathematical Reasoning

Thang Luong, Dawsen Hwang, Hoang Nguyen, Golnaz Ghiasi, Yuri Chervonyi, Insuk Seo, Junsu Kim, G. Bingham, Jonathan Lee, Swaroop Mishra, A. Zhai, C. Hu, H. Michalewski, Jimin Kim, J. Ahn, Junhwi Bae, Xingyou Song, Trieu H. Trinh, Quoc V. Le, Junehyuk Jung
November 3, 2025Conference on Empirical Methods in Natural Language Processing59 citations

59

Citations

14

Influential Citations

Conference on Empirical Methods in Natural Language Processing

Venue

2025

Year

Abstract

Finding the right north-star metrics is highly critical for advancing the mathematical reasoning capabilities of foundation models, especially given that existing evaluations are either too easy or only focus on getting correct short answers. To address these issues, we present IMO-Bench, a suite of advanced reasoning benchmarks, vetted by a panel of top specialists and that specifically targets the level of the International Mathematical Olympiad (IMO), the most prestigious venue for young mathematicians. IMO-AnswerBench first tests models on 400 diverse Olympiad problems with verifiable short answers. IMO-Proof Bench is the next-level evaluation for proof-writing capabilities, which includes both basic and advanced IMO level problems as well as detailed grading guidelines to facilitate automatic grading. These benchmarks played a crucial role in our historic achievement of the gold-level performance at IMO 2025 with Gemini Deep Think (Luong and Lockhart, 2025). Our model achieved 80.0% on IMO-AnswerBench and 65.7% on the advanced IMO-Proof Bench, surpassing the best non-Gemini models by large margins of 6.9% and 42.4% respectively. We also showed that autograders built with Gemini reasoning correlate well with human evaluations and construct IMO-GradingBench, with 1000 human gradings on proofs, to enable further progress in automatic evaluation of long-form answers. We hope that IMO-Bench will help the community towards advancing robust mathematical reasoning and release it at https://imobench.github.io/.

Analysis

Why This Paper Matters

This paper addresses a critical gap in evaluating mathematical reasoning for foundation models: existing benchmarks are either too simplistic (e.g., GSM8K, MATH) or only measure short-answer correctness, failing to capture the depth of proof-based reasoning required at the highest levels. By introducing IMO-Bench, the authors provide a suite of problems vetted by top specialists that directly target the International Mathematical Olympiad (IMO) standard—the most prestigious venue for young mathematicians. This is significant because advancing AI to gold-level IMO performance has been a long-standing challenge, and the paper reports that Gemini Deep Think achieves this milestone, surpassing all non-Gemini models by large margins. The work also tackles the difficult problem of automatic grading of proofs, constructing IMO-GradingBench with 1000 human gradings to validate autograders. This combination of rigorous benchmarks and validated evaluation tools sets a new standard for the field.

Technical Contributions

  • IMO-AnswerBench: A collection of 400 diverse Olympiad problems with verifiable short answers, designed to test factual and computational reasoning without the complexity of full proofs.
  • IMO-Proof Bench: A two-tier evaluation for proof-writing: basic and advanced IMO-level problems, each accompanied by detailed grading guidelines to enable automatic grading.
  • Automatic Grading with Gemini: The authors build autograders using Gemini reasoning and show they correlate well with human evaluations, reducing the need for manual grading of long-form answers.
  • IMO-GradingBench: A dataset of 1000 human gradings on proofs, released to facilitate further progress in automatic evaluation.
  • Gold-Level IMO Performance: The paper demonstrates that Gemini Deep Think achieves 80.0% on IMO-AnswerBench and 65.7% on advanced IMO-Proof Bench, outperforming the best non-Gemini models by 6.9% and 42.4% respectively.

Results

The key quantitative results are: Gemini Deep Think scores 80.0% on IMO-AnswerBench and 65.7% on advanced IMO-Proof Bench. Compared to the best non-Gemini models, this represents a 6.9% absolute improvement on the answer benchmark and a 42.4% absolute improvement on the proof benchmark. The autograders built with Gemini reasoning show strong correlation with human evaluations, though exact correlation coefficients are not provided in the abstract. The paper also reports that these benchmarks were instrumental in achieving gold-level performance at IMO 2025, a historic milestone.

Significance

This work has broad implications for AI research and education. First, it provides a rigorous, expert-vetted benchmark that can drive progress in mathematical reasoning, a key capability for scientific discovery and engineering. Second, the demonstration of gold-level IMO performance shows that foundation models are approaching human-level expertise in formal reasoning domains. Third, the release of IMO-GradingBench and the autograding methodology addresses a practical bottleneck in evaluating long-form answers, which is relevant beyond mathematics (e.g., in code generation, legal reasoning). The benchmarks and tools are publicly released, enabling the community to build on this work and further advance robust mathematical reasoning.