MixEval
Freea ground-truth-based dynamic benchmark derived from off-the-shelf benchmark mixtures, which evaluates LLMs with a highly capable model ranking (i.e., 0.96 correlation with Chatbot Arena) while running locally and quickly (6% the time and cost of running MMLU).
About MixEval
MixEval is a new paradigm for LLM evaluation that strategically mixes off-the-shelf benchmarks to bridge real-world user queries and ground-truth-based evaluation. It achieves a 0.96 ranking correlation with Chatbot Arena while being fast, cheap (6% time and cost of MMLU), and reproducible. The benchmark features dynamic monthly updates to mitigate contamination, and includes a harder version, MixEval-Hard. Accepted at NeurIPS 2024, MixEval provides a leaderboard and open-source evaluation suite that can be run locally.
Key Features
Pros & Cons
- Highest correlation with human judgments among leading benchmarks
- Cost-effective and quick evaluation
- Dynamic updates prevent benchmark contamination
- Reproducible and ground-truth-based
- Requires local setup for evaluation
- Base version is text-only (MixEval-X covers multimodal)