
We're excited to team up with Kaggle for a brand new challenge! Running through October 11, the...
We're excited to team up with Kaggle for a brand new challenge!
Running through October 11, the Kaggle Benchmarking Challenge asks you to build a benchmark, run it against real models, and tell us what you found.
Today's models reason, write code, use tools, and hold multi-turn conversations, and their capabilities will likely change by tomorrow. Public leaderboards only tell you so much, so why not test them against the work you actually do? Kaggle Benchmarks lets you build your own evaluations and run them across a suite of models to see how they really stack up.
So pick something you've always wondered about, measure it, and show your work. There's a $2,500 prize pool up for grabs across five winners!
Read on to learn more.
{% card %}
Your mandate is to build a benchmark on Kaggle and write about what you learned.
Your post should cover:
The scope is wide open. Benchmark multi-step reasoning, code generation, tool use, image recognition, instruction following, or that one weird failure mode you keep running into. The best benchmarks usually come from a specific itch rather than a general one.
<center> {% cta https://dev.to/new?prefill=---%0Atitle%3A%20%0Apublished%3A%20%0Atags%3A%20devchallenge%2C%20kagglechallenge%2C%20ai%2C%20machinelearning%0A---%0A%0A%2AThis%20is%20a%20submission%20for%20the%20%5BKaggle%20Benchmarking%20Challenge%5D%28https%3A%2F%2Fdev.to%2Fchallenges%2Fkaggle-2026-09-23%29%2A%0A%0A%23%23%20What%20I%20Benchmarked%0A%3C%21--%20What%20task%28s%29%20did%20you%20run%3F%20Tell%20us%20what%20capability%20or%20behavior%20you%20set%20out%20to%20measure%20and%20why%20it%20interested%20you.%20--%3E%0A%0A%23%23%20Models%20Tested%0A%3C%21--%20Which%20models%20did%20you%20run%20your%20benchmark%20against%2C%20and%20why%20did%20you%20pick%20them%3F%20--%3E%0A%0A%23%23%20Findings%0A%3C%21--%20Share%20your%20results%20and%20the%20main%20insights.%20What%20surprised%20you%3F%20What%20would%20you%20measure%20next%3F%20--%3E%0A%0A%23%23%20My%20Benchmark%0A%3C%21--%20Required%3A%20share%20a%20link%20to%20your%20benchmark%20on%20Kaggle%20so%20judges%20and%20readers%20can%20check%20it%20out.%20--%3E%0A%0A%3C%21--%20Don%27t%20forget%20to%20add%20a%20cover%20image%20and%20include%20any%20other%20appropriate%20tags%20for%20your%20post.%20--%3E%0A%0A%3C%21--%20Thanks%20for%20participating%21%20--%3E %} Kaggle Benchmarking Challenge Submission Template {% endcta %} </center>
The most compelling submissions will go beyond reporting numbers. Tell us what the results actually mean and what they changed about how you think about these models.
{% endcard %}
All submissions will be evaluated on:
Note: Submissions must include a link to the benchmark on Kaggle to be eligible.
Five winners will each receive:
All participants with a valid submission will receive a completion badge on their DEV profile.
New to Kaggle Benchmarks? Here's how it works:
Prefer your own setup? You can now create tasks from your local development environment using the Kaggle CLI.
Helpful resources:
Read more from the Kaggle team right here on DEV:
{% embed https://dev.to/googleai/introducing-community-benchmarks-on-kaggle-35nc %}
{% embed https://dev.to/googleai/kaggle-is-making-ai-benchmark-creation-effortless-1g7n %}
{% card %}
Build your benchmark on Kaggle, then publish a post on DEV using the submission template provided above and the required challenge tag: #kagglechallenge.
One submission per participant, so make it count!
Please review our judging criteria, rules, guidelines, and FAQ page before submitting so you understand our participation guidelines and official contest rules such as eligibility requirements.
{% endcard %}
We can't wait to see what you measure. Questions about the challenge? Drop them in the comments below.
Good luck and happy benchmarking!
eventsinyourcityOn Wednesday, September 30, Anthropic and Google Cloud are co-hosting a hands-on developer workshop...
opensourceFor decades, platforms like Claris FileMaker, Microsoft Access, and 4D enabled businesses to build...
aiThis video (the script, the voice timings, the source code, the storyboard, the briefs the subagents...
devchallengeWe are thrilled to announce the winners of DEV's Big Summer Bug Smash powered by Sentry! This...
gemmaGemma 4 E2B, E4B, 12B and a 26B-A4B fp8 build read by their label probabilities with vLLM on one TPU v6e chip, checked against the same read on an NVIDIA L4 and against Jev 1.13.0's published results. What fits one chip, how to read labels when vLLM on TPU returns only the top 32 log-probabilities, speed, cost, and why no 31B loads today.
aiCoding agents in your terminal can scaffold microservices in seconds. But what happens when your...
Workflows from the Neura Market marketplace related to this Perplexity resource