MixEval logo

MixEval

Free

a ground-truth-based dynamic benchmark derived from off-the-shelf benchmark mixtures, which evaluates LLMs with a highly capable model ranking (i.e., 0.96 correlation with Chatbot Arena) while running locally and quickly (6% the time and cost of running MMLU).

FreeFree tier
Inputs: text
Type
Open Source

About MixEval

MixEval is a new paradigm for LLM evaluation that strategically mixes off-the-shelf benchmarks to bridge real-world user queries and ground-truth-based evaluation. It achieves a 0.96 ranking correlation with Chatbot Arena while being fast, cheap (6% time and cost of MMLU), and reproducible. The benchmark features dynamic monthly updates to mitigate contamination, and includes a harder version, MixEval-Hard. Accepted at NeurIPS 2024, MixEval provides a leaderboard and open-source evaluation suite that can be run locally.

Key Features

High correlation with Chatbot Arena (0.96)
Fast and cheap evaluation (6% cost of MMLU)
Dynamic monthly updates to prevent contamination
Ground-truth-based scoring
MixEval-Hard for harder evaluation
Open-source and locally runnable

Pros & Cons

Pros
  • Highest correlation with human judgments among leading benchmarks
  • Cost-effective and quick evaluation
  • Dynamic updates prevent benchmark contamination
  • Reproducible and ground-truth-based
Cons
  • Requires local setup for evaluation
  • Base version is text-only (MixEval-X covers multimodal)

Best For

Evaluating LLMs for research and developmentComparing model performance reliablyBenchmarking in academic studiesReal-world query simulation

FAQ

What is MixEval?
MixEval is a ground-truth-based dynamic benchmark derived from off-the-shelf benchmark mixtures, designed to evaluate LLMs with high correlation to Chatbot Arena while being fast and cheap.
How often is MixEval updated?
MixEval is updated monthly to mitigate contamination risk.
What is the cost of running MixEval?
It costs approximately 6% of the time and cost of running MMLU.
What is MixEval-Hard?
MixEval-Hard is a harder version of MixEval that offers more room for model improvement.