MixEval logo

MixEval

Free

A reliable click-and-go evaluation suite compatible with both open-source and proprietary models, supporting MixEval and other benchmarks.

FreeFree tier
Type
Open Source

About MixEval

MixEval is a ground-truth-based dynamic benchmark for evaluating large language models (LLMs). It derives queries from off-the-shelf benchmark mixtures and achieves a 0.96 correlation with Chatbot Arena Elo while running locally and quickly—at 6% the time and cost of MMLU. The evaluation suite supports both open-source and proprietary models, offering click-and-go model response generation and score computation. MixEval includes two benchmarks (MixEval and MixEval-Hard) with free-form and multiple-choice splits, and its queries are automatically updated monthly to avoid contamination. The suite uses a model parser (GPT-3.5-Turbo or open-source) for stable scoring, and allows easy registration of custom models and benchmark data.

Key Features

Dynamic benchmark derived from off-the-shelf benchmark mixtures
0.96 correlation with Chatbot Arena Elo
Runs locally with 6% time and cost of MMLU
Monthly automatic updates to prevent data contamination
Includes MixEval and MixEval-Hard with free-form and multiple-choice splits
Supports both open-source and proprietary models
Click-and-go model response generation and score computation
Uses GPT-3.5-Turbo or open-source model as stable parser
Easy registration of custom models and benchmark data

Pros & Cons

Pros
  • High correlation with human preference (Chatbot Arena)
  • Very cost-effective: 6% time and cost of MMLU
  • Dynamic updates reduce contamination risk
  • Easy to use with click-and-go setup
  • Supports both open-source and proprietary models
  • Provides stable model parser for scoring
Cons
  • Requires external parser (GPT-3.5 or open-source model) for scoring, which may incur cost
  • Benchmark derived from existing benchmarks may inherit their biases
  • Currently only supports text-to-text evaluation (separate MixEval-X for any-to-any)
  • Setup requires cloning repo and managing dependencies

Best For

Evaluating and ranking LLMs in researchComparing model performance against human preferencesCost-effective alternative to Chatbot Arena for model benchmarkingQuick iterative testing of model checkpointsAcademic benchmarking and leaderboard creation

FAQ

What is MixEval?
MixEval is a ground-truth-based dynamic benchmark that evaluates LLMs with high correlation to Chatbot Arena (0.96). It is derived from off-the-shelf benchmark mixtures and runs locally at a fraction of the cost and time of MMLU, with monthly updates to avoid contamination.
How does MixEval compare to other benchmarks?
MixEval achieves a 0.96 correlation with Chatbot Arena Elo, while being significantly more cost-effective and faster than MMLU (6% time and cost). It offers monthly dynamic updates, unlike static benchmarks.