OlympicArena logo

OlympicArena

Free

a benchmark for evaluating AI models across multiple academic disciplines like math, physics, chemistry, biology, and more.

FreeFree tier
Inputs: text, imageOutputs: text
Type
Open Source

About OlympicArena

OlympicArena is a comprehensive, high-challenging, and rigorously curated benchmark designed to assess advanced AI capabilities across a broad spectrum of Olympic-level challenges. It comprises 11,163 problems from 62 distinct Olympic competitions, structured with 13 answer types and spanning seven core disciplines: mathematics, physics, chemistry, biology, geography, astronomy, and computer science, encompassing 34 specialized branches. The benchmark focuses on Olympic-level problems and covers 8 types of logical reasoning abilities and 5 types of visual reasoning abilities. It employs an instance-level leakage detection metric to ensure benchmark integrity. Evaluations are conducted from both answer-level and process-level perspectives, with fine-grained analyses on different types of cognitive reasoning. The project includes a leaderboard for LLMs and LMMs (both open-source and proprietary) using zero-shot prompts, and all data and code are open-sourced.

Key Features

Comprehensive: 11,163 problems from 62 Olympic competitions across 7 disciplines and 34 branches
High-challenging: Olympic-level problems covering 8 logical and 5 visual reasoning abilities
Rigorous: Uses instance-level leakage detection to validate benchmark effectiveness
Fine-grained evaluation: Answer-level and process-level metrics with cognitive reasoning analysis
Multimodal support: Evaluates both text-only LLMs and vision-language LMMs
Open-source: Code, dataset on Hugging Face, and leaderboard publicly available

Pros & Cons

Pros
  • Extensive coverage of diverse Olympic-level problems across seven STEM disciplines
  • Both answer-level and process-level evaluation for nuanced insights
  • Open-source dataset and code promote reproducibility and further research
  • Includes rigorous leakage detection to ensure benchmark validity
  • Supports both text-only and multimodal models for broader capability assessment
Cons
  • Current evaluation uses zero-shot prompts only, no few-shot or chain-of-thought comparison
  • Limited to specific Olympic competition problems, not covering all possible reasoning domains
  • Answer extraction relies on rule-based matching which may not handle all free-form responses

Best For

Evaluating AI models' cognitive reasoning in STEM disciplinesComparing performance of LLMs vs LMMs on complex multi-step problemsBenchmarking progress toward superintelligent AI in scientific problem-solvingAssessing leakage in pre-training corpora via instance-level detection