MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering logo

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Free

Benchmarking AI agents on real-world ML engineering through Kaggle competitions

FreeFree tier
Type
Open Source

About MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

MLE-bench is a benchmark designed to evaluate the performance of AI agents on machine learning engineering tasks. It curates 75 ML engineering-related competitions from Kaggle, covering diverse real-world challenges such as training models, preparing datasets, and running experiments. The benchmark provides human baselines derived from Kaggle leaderboards and assesses frontier language models using open-source agent scaffolds. In the best-performing setup—OpenAI's o1-preview with AIDE scaffolding—agents achieve at least a Kaggle bronze medal in 16.9% of competitions. The benchmark code is open-source, facilitating further research into AI agents' ML engineering capabilities.

Key Features

Based on 75 curated ML engineering competitions from Kaggle
Tests real-world skills: training models, preparing datasets, running experiments
Provides human baselines from Kaggle leaderboards
Evaluates frontier language models with open-source agent scaffolds
Investigates resource scaling and pre-training contamination effects
Open-source code available for further research

Pros & Cons

Pros
  • Covers diverse, real-world ML engineering tasks from Kaggle
  • Provides human performance baselines for context
  • Open-source implementation encourages reproducibility and extensions
  • Includes analysis of resource scaling and data contamination
Cons
  • Focused primarily on Kaggle-style competitions, which may not reflect all ML engineering contexts
  • High computational cost required to run evaluations on frontier models
  • Agent performance may not generalize to non-Kaggle or custom ML engineering workflows

Best For

Evaluating AI agent performance in machine learning engineeringResearch on AI capabilities for end-to-end ML tasksComparing agent scaffolds and language models on realistic ML challengesStudying scaling behavior and contamination in AI benchmarks

FAQ

What is MLE-bench?
MLE-bench is a benchmark for evaluating how well AI agents perform at machine learning engineering, using 75 competitions from Kaggle.
How many competitions are included in MLE-bench?
The benchmark curates 75 ML engineering-related competitions from Kaggle.
What was the best performance achieved by AI agents?
The best-performing setup, OpenAI's o1-preview with AIDE scaffolding, achieved at least a Kaggle bronze medal in 16.9% of competitions.
Are human baselines available?
Yes, human baselines are established using Kaggle's publicly available leaderboards for each competition.
Is the benchmark code open-source?
Yes, the benchmark code is open-source and available on GitHub.