PreprintConference on Empirical Methods in Natural Language Processing2025
Towards Robust Mathematical Reasoning
Thang Luong, Dawsen Hwang, Hoang Nguyen, et al.
Introduces IMO-Bench, a suite of Olympiad-level benchmarks for robust mathematical reasoning, enabling gold-level IMO performance via Gemini Deep Think.
59Nov 3, 2025ReasoningBenchmarks
arXiv