LLM Stats logo

LLM Stats

Free

Compare AI models across benchmarks, pricing, speed, and context window.

FreeFree tier
Type
Open Source

About LLM Stats

LLM Stats provides an independent AI leaderboard ranking over 300 top AI models across intelligence, speed, and price. The site aggregates public benchmarks and live API metrics into a composite LLM Stats Score, allowing users to compare models like GPT, Claude, Gemini, Llama, DeepSeek, and more. It includes detailed filters, per-axis leaders (e.g., cheapest in top 10, fastest output, longest context), and continuously updated rankings.

Key Features

Aggregates GPQA, SWE-Bench Verified, coding-arena performance and pricing into one LLM Stats Score
Ranks over 300 models including GPT, Claude, Gemini, Llama, DeepSeek, Qwen, Mistral, GLM
Continuously updated from public benchmarks and live API metrics
Advanced filters for time range (90d, 30d), open-source vs proprietary
Per-axis highlights: fastest output, cheapest in top 10, longest context window
Includes model parameters, context length, speed (tokens/sec), and pricing ($/M tokens)

Pros & Cons

Pros
  • Independent, unbiased rankings based on publicly available data
  • Covers a broad set of models from multiple providers
  • Updates continuously as new benchmarks and pricing changes occur
  • Transparent methodology with composite score explained
  • Free to use with no login required
Cons
  • Rankings rely on public benchmarks which may not reflect all real-world use cases
  • Some models may have missing data (e.g., unreleased previews show dashes for speed/pricing)
  • Does not include all possible AI models, only those meeting inclusion criteria

Best For

Selecting the best AI model for specific tasks like reasoning, coding, or agentic workflowsComparing costs and performance trade-offs among frontier modelsTracking the evolving AI landscape for research or procurement decisionsIdentifying leading open-weight models for self-hosting

FAQ

Which AI model currently ranks #1 on the LLM Leaderboard?
Claude Mythos Preview currently leads on GPQA Diamond (94.6% gpqa), the most discriminating reasoning benchmark. The overall LLM Stats Score ranking is shown in the table.
What is the best AI model right now?
Best depends on optimization goal. For frontier reasoning, Claude Mythos Preview leads on GPQA. For coding agents, the current leader is the strongest in head-to-head coding-arena. For low cost at frontier quality, Grok 4.5 is the cheapest in the top 10 at $2.00/M tok.
What are the best LLMs in 2026?
Leading LLMs in 2026 include Claude Mythos Preview, GPT-5 family, Claude Opus and Sonnet, Gemini 3 Pro, Grok 4, DeepSeek V3/R1, and Z.AI GLM-5. Open-weights leaders include Llama, Qwen, and DeepSeek.
What is the cheapest AI model in the top 10?
Grok 4.5 is the cheapest model in the top 10 by GPQA Diamond, at $2.00 per million tokens.