OlympicArena
Freea benchmark for evaluating AI models across multiple academic disciplines like math, physics, chemistry, biology, and more.
About OlympicArena
OlympicArena is a comprehensive, high-challenging, and rigorously curated benchmark designed to assess advanced AI capabilities across a broad spectrum of Olympic-level challenges. It comprises 11,163 problems from 62 distinct Olympic competitions, structured with 13 answer types and spanning seven core disciplines: mathematics, physics, chemistry, biology, geography, astronomy, and computer science, encompassing 34 specialized branches. The benchmark focuses on Olympic-level problems and covers 8 types of logical reasoning abilities and 5 types of visual reasoning abilities. It employs an instance-level leakage detection metric to ensure benchmark integrity. Evaluations are conducted from both answer-level and process-level perspectives, with fine-grained analyses on different types of cognitive reasoning. The project includes a leaderboard for LLMs and LMMs (both open-source and proprietary) using zero-shot prompts, and all data and code are open-sourced.
Key Features
Pros & Cons
- Extensive coverage of diverse Olympic-level problems across seven STEM disciplines
- Both answer-level and process-level evaluation for nuanced insights
- Open-source dataset and code promote reproducibility and further research
- Includes rigorous leakage detection to ensure benchmark validity
- Supports both text-only and multimodal models for broader capability assessment
- Current evaluation uses zero-shot prompts only, no few-shot or chain-of-thought comparison
- Limited to specific Olympic competition problems, not covering all possible reasoning domains
- Answer extraction relies on rule-based matching which may not handle all free-form responses