Compare AI model performance across 48+ benchmarks. Our composite index aggregates coding, math, reasoning, and language scores into a single intelligence ranking.
At-a-glance rankings across the three dimensions that matter most.
Neura Intelligence Index; Higher is better
Max input tokens; Higher is better
USD per 1M output tokens; Lower is better
Ranked by the Neura Intelligence Index — a weighted composite of 48 benchmarks across 8 categories.
Find the sweet spot — models in the top-left quadrant offer the best value.
Top performers in each benchmark category.
Who leads on each individual benchmark — click any card to see full results.
Different tasks need different strengths. These indices re-weight our benchmarks for specific workflows.
Best for software development, code generation, and debugging
Best for scientific research, data analysis, and complex reasoning
Best for writing, editing, summarization, and creative tasks
Highest intelligence per dollar — the cost-efficiency sweet spot
The Neura Intelligence Index is a composite score (0-100) computed from 48+ individual benchmarks spanning 8 categories: coding, math, reasoning, general knowledge, language, multimodal, safety, and agentic tasks.
Not all models have scores on all benchmarks. The confidence indicator reflects benchmark coverage: high (>70% of benchmarks), medium (40-70%), or low (<40%). Weights are renormalized across available categories so models aren't penalized for missing data.
Scores are aggregated from official model cards, Papers With Code, HuggingFace Open LLM Leaderboard, LiveBench, and LMSYS Chatbot Arena. Each score includes a verification status (official, self-reported, or aggregated).
The Neura Intelligence Index is a composite score (0-100) that aggregates AI model performance across 15+ benchmarks spanning coding, math, reasoning, general knowledge, language, multimodal, safety, and agentic tasks. It uses min-max normalization and weighted category averaging to produce a single comparable score.
Benchmark scores are synced daily via an automated pipeline that aggregates data from official model cards, Papers With Code, HuggingFace Open LLM Leaderboard, LiveBench, and LMSYS Chatbot Arena.
Confidence reflects how many benchmarks a model has been tested on relative to the total. High means >70% benchmark coverage, medium is 40-70%, and low is <40%. Models with low coverage may have composite scores that shift as more benchmarks are added.
Category weights reflect real-world demand: Coding and Reasoning each get 20%, General and Math each get 15%, Language and Multimodal each get 10%, and Safety and Agent each get 5%. Weights are renormalized across available categories so models are not penalized for missing data.
| # | Model | Coverage | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
1— | Claude Opus 5.5 Anthropic🇺🇸1w | 99.9 | 99.9 | — | — | — | — | — | $20 | — | — | low |
2— | X Grok 4.6 xAI🇺🇸1mo | 99.6 | — | — | — | — | — | — | $6 | — | — | low |
3— | Claude Sonnet 5.5 Anthropic🇺🇸3d | 99.0 | 99.0 | — | — | — | — | — | $10 | — | — | low |
4— | GPT-Image-1 OpenAI🇺🇸1y | 98.7 | — | — | — | — | — | — | — | — | — | low |
5— | Claude Fable 5 Anthropic🇺🇸3mo | 97.0 | 97.5 | — | 96.4 | — | — | — | $50 | — | — | low |
6— | Claude Fable 5.1 Anthropic🇺🇸4w | 96.6 | — | — | 96.6 | — | — | — | $50 | — | — | low |
7— | M Muse Spark 1.3 Meta🇺🇸4w | 96.0 | — | — | — | 96.0 | — | — | $4 | — | — | low |
8— | Claude Opus 5 Anthropic🇺🇸2mo | 95.2 | — | — | 96.5 | — | — | — | $25 | — | — | low |
9— | Z GLM-5.3 Zhipu AI🇨🇳1moOSS | 94.8 | — | — | 95.4 | — | — | — | $4 | — | — | low |
10— | GPT-6 Astra OpenAI🇺🇸3w | 94.2 | — | — | 94.1 | — | — | 96.3 | $50 | — | — | medium |
11— | DeepSeek-V4-Pro-0813 DeepSeek🇨🇳1moOSS | 93.8 | — | — | 93.9 | — | — | — | $3 | — | — | low |
12— | Claude Mythos Preview Anthropic🇺🇸 | 91.4 | 96.7 | — | 93.8 | 83.8 | — | 90.2 | — | — | — | medium |
13— | Gemini 3.7 Flash Google🇺🇸1mo | 90.7 | — | — | — | 95.5 | — | 83.5 | $4 | — | — | low |
14— | Z GLM-5.3-Flash Zhipu AI🇨🇳1moOSS | 89.5 | — | — | 90.1 | — | — | 84.7 | $0.50 | — | — | low |
15— | Sora OpenAI🇺🇸1y | 89.4 | — | — | — | — | — | — | — | — | — | low |
16— | DeepSeek-V4-Flash-0731 DeepSeek🇨🇳2moOSS | 89.1 | — | — | — | — | — | — | $0.18 | — | — | low |
17— | Claude Opus 4.8 Anthropic🇺🇸4mo | 89.1 | 89.4 | — | 91.3 | — | — | 89.3 | $25 | — | — | high |
18— | M Kimi K3 Moonshot AI🇨🇳2moOSS | 88.2 | — | — | 90.4 | — | — | 85.8 | $14 | — | — | high |
19— | Veo 2 Google🇺🇸1y | 88.2 | — | — | — | — | — | — | — | — | — | low |
20— | GPT-5.6 Sol OpenAI🇺🇸2mo | 87.4 | 74.4 | 98.2 | 91.0 | 93.3 | — | 86.3 | $30 | — | — | medium |
21— | Q Qwen3.8 Max Alibaba Cloud / Qwen Team🇨🇳2moOSS | 86.3 | 83.3 | — | 81.6 | 94.0 | — | 87.4 | $5 | — | — | medium |
22— | GPT-5.6 Terra OpenAI🇺🇸2mo | 84.9 | 70.4 | 97.4 | 89.6 | 92.4 | — | 82.3 | $12 | — | — | medium |
23— | X Grok-4 Heavy xAI🇺🇸 | 84.7 | — | 84.3 | 84.9 | — | — | — | — | — | — | low |
24— | GPT-5.1 High OpenAI🇺🇸10mo | 84.3 | — | 83.8 | 84.5 | — | — | — | — | — | — | low |
25— | Claude Opus 4.7 Anthropic🇺🇸5mo | 84.2 | 82.3 | — | 90.1 | 80.6 | — | 87.2 | $25 | — | — | high |
26— | Claude Sonnet 5 Anthropic🇺🇸3mo | 83.5 | 79.1 | — | 92.0 | — | — | 82.8 | $10 | — | — | medium |
27— | T Hy4 preview Tencent🇨🇳1moOSS | 83.2 | 77.8 | — | 89.0 | — | — | — | — | — | — | medium |
28— | GPT-5.1 Medium OpenAI🇺🇸10mo | 82.4 | — | 82.4 | — | — | — | — | $10 | — | — | low |
29— | X Grok 4.5 xAI🇺🇸2mo | 82.2 | 74.7 | — | 89.7 | — | — | — | $6 | — | — | low |
30— | M Muse Spark 1.1 Meta🇺🇸2mo | 81.3 | 63.6 | — | 95.2 | — | — | 83.0 | $4 | — | — | medium |
31— | DALL-E 3 OpenAI🇺🇸3y | 81.0 | — | — | — | — | — | — | — | — | — | low |
32— | GPT-5 High OpenAI🇺🇸1y | 80.8 | — | 77.2 | 83.6 | — | — | — | — | — | — | low |
33— | B Seed 2.1 Pro ByteDance🇨🇳3mo | 80.8 | 72.4 | — | 90.5 | — | — | 82.5 | — | — | — | high |
34— | Claude Opus 4.6 Anthropic🇺🇸7mo | 80.2 | 82.4 | 84.1 | 85.2 | 81.1 | — | 66.0 | $25 | — | — | high |
35— | GPT-5.1 Codex High OpenAI🇺🇸10mo | 80.2 | — | 80.2 | — | — | — | — | — | — | — | low |
36— | B Seed 2.0 Pro ByteDance🇨🇳7mo | 80.0 | 74.8 | 82.3 | 85.5 | — | — | — | $3 | — | — | medium |
37— | ChatGPT-4o Latest OpenAI🇺🇸2y | 79.1 | — | — | 79.1 | — | — | — | — | — | — | low |
38— | Gemini 3.8 Flash Google🇺🇸4w | 78.8 | — | — | — | — | — | 78.8 | $4 | — | — | low |
39— | Claude Sonnet 4.5 Anthropic🇺🇸1y | 78.7 | 92.8 | 64.7 | 78.3 | 72.9 | — | — | $15 | — | — | medium |
40— | GPT-5.5 Pro OpenAI🇺🇸5mo | 78.6 | — | 57.5 | 91.8 | — | — | — | — | — | — | low |
41— | Claude Code Anthropic🇺🇸1y | 78.6 | 66.8 | — | — | — | — | — | — | — | — | low |
42— | B Seed 2.1 Turbo ByteDance🇨🇳3mo | 78.4 | 70.4 | — | 89.4 | — | — | 79.0 | — | — | — | high |
43— | X MiMo-V2-Pro Xiaomi🇨🇳6mo | 77.6 | 77.6 | — | — | — | — | — | — | — | — | low |
44— | X Grok-3 xAI🇺🇸1y | 77.5 | — | 75.3 | 80.0 | — | — | 75.9 | $15 | — | — | low |
45— | GPT-5 Medium OpenAI🇺🇸1y | 77.5 | — | 68.0 | 84.5 | — | — | — | — | — | — | low |
46— | Q Qwen3.7 Max Alibaba Cloud / Qwen Team🇨🇳4mo | 76.9 | 76.5 | — | 79.4 | 76.9 | — | — | $4 | — | — | medium |
47— | Z GLM-5V-Turbo Zhipu AI🇨🇳6mo | 76.5 | — | — | — | — | — | — | — | — | — | low |
48— | ERNIE 5.0 Baidu🇨🇳8mo | 76.3 | — | 64.7 | 72.8 | 92.5 | — | — | — | — | — | medium |
49— | GPT-5.2 Pro OpenAI🇺🇸9mo | 76.2 | — | 84.3 | 70.8 | — | — | — | — | — | — | medium |
50— | X Grok-3 Mini xAI🇺🇸1y | 75.8 | — | 71.3 | 79.1 | — | — | — | — | — | — | low |
51— | GPT-5.1 OpenAI🇺🇸10mo | 75.8 | 74.4 | 57.1 | 84.5 | — | — | 89.0 | $10 | — | — | medium |
52— | GPT-5.1 Instant OpenAI🇺🇸10mo | 75.8 | 74.4 | 57.1 | 84.5 | — | — | 89.0 | $10 | — | — | medium |
53— | GPT-5.1 Thinking OpenAI🇺🇸10mo | 75.8 | 74.4 | 57.1 | 84.5 | — | — | 89.0 | — | — | — | medium |
54— | Claude Sonnet 4.6 Anthropic🇺🇸7mo | 75.7 | 80.4 | — | 79.0 | 73.6 | — | 71.0 | $15 | — | — | high |
55— | Q Qwen3.8 Flash Alibaba Cloud / Qwen Team🇨🇳1mo | 75.3 | 67.2 | — | 73.4 | — | — | 86.6 | $0.47 | — | — | medium |
56— | Q Qwen3.8-Flash-Next Alibaba Cloud / Qwen Team🇨🇳1moOSS | 75.3 | 67.2 | — | 73.4 | — | — | 86.6 | — | — | — | medium |
57— | o1 OpenAI🇺🇸1y | 75.2 | 48.4 | 94.2 | 79.3 | 86.7 | — | — | $60 | — | — | high |
58— | M Kimi K2-Thinking-0905 Moonshot AI🇨🇳1yOSS | 75.1 | 69.5 | 84.3 | 82.6 | — | — | — | — | — | — | medium |
59— | Claude Opus 4 Anthropic🇺🇸1y | 74.8 | 58.0 | 75.5 | 79.3 | 96.0 | — | 66.8 | $75 | — | — | high |
60— | B Seed 2.0 Lite ByteDance🇨🇳7mo | 74.7 | 68.5 | 74.8 | 80.7 | — | — | — | — | — | — | low |
61— | Z GLM-5.2 Zhipu AI🇨🇳3moOSS | 74.6 | 65.8 | — | 88.7 | — | — | — | $2 | — | — | medium |
62— | DeepSeek-V4.1-Flash DeepSeek🇨🇳3wOSS | 74.0 | — | — | 74.0 | — | — | — | $0.66 | — | — | low |
63— | GPT-5.6 Luna OpenAI🇺🇸2mo | 73.6 | 68.0 | 95.5 | 89.0 | 38.0 | — | 77.5 | $1 | — | — | medium |
64— | DeepSeek-V4-Pro-Max DeepSeek🇨🇳5moOSS | 73.1 | 60.9 | — | 84.1 | 78.1 | — | — | $3 | — | — | high |
65— | Midjourney v6.1 Midjourney🇺🇸2y | 73.0 | — | — | — | — | — | — | — | — | — | low |
66— | T Hy3 Tencent🇨🇳2moOSS | 72.9 | 63.6 | — | 87.1 | — | — | — | $0.58 | — | — | medium |
67— | S Step-3.5-Flash StepFun🇨🇳8moOSS | 72.8 | 70.5 | 81.0 | — | — | — | — | $0.40 | — | — | low |
68— | Z GLM-5 Zhipu AI🇨🇳7moOSS | 72.7 | 77.2 | — | — | — | — | — | $3 | — | — | low |
69— | Gemini 3.1 Pro Google🇺🇸7mo | 72.4 | 71.2 | — | 87.9 | 51.8 | — | 81.9 | $12 | — | — | high |
70— | GPT-5.2 OpenAI🇺🇸9mo | 72.4 | 81.1 | 71.4 | 68.4 | 74.6 | — | 80.3 | $14 | — | — | high |
71— | M Kimi K2.6 Moonshot AI🇨🇳5moOSS | 72.3 | 72.7 | — | 73.3 | — | — | 80.4 | $4 | — | — | high |
72— | L EXAONE 4.5 33B LG AI Research🇰🇷5mo | 72.0 | — | 74.6 | 73.8 | — | — | 64.5 | — | — | — | medium |
73— | Gemini 3 Pro Google🇺🇸10mo | 72.0 | 74.2 | 84.3 | 65.7 | 64.1 | — | 73.2 | — | — | — | high |
74— | Claude Opus 4.5 Anthropic🇺🇸10mo | 71.9 | 82.6 | — | 61.3 | 78.5 | — | — | — | — | — | medium |
75— | M Kimi K2.5 Moonshot AI🇨🇳8moOSS | 71.9 | 57.2 | 79.3 | 84.1 | — | — | 67.5 | — | — | — | high |
76— | X MiMo-V2-Omni Xiaomi🇨🇳6mo | 71.3 | 71.3 | — | — | — | — | — | — | — | — | low |
77— | T Inkling Thinking Machines Lab🇺🇸2moOSS | 71.3 | 76.9 | — | — | — | — | 62.2 | $4 | — | — | medium |
78— | o1-pro OpenAI🇺🇸1y | 71.2 | — | — | 71.2 | — | — | — | — | — | — | low |
79— | GPT OSS 20B High OpenAI🇺🇸1yOSS | 71.2 | — | 82.8 | 62.5 | — | — | — | — | — | — | low |
80— | Q Qwen3.7-Plus Alibaba Cloud / Qwen Team🇨🇳4mo | 71.2 | 69.0 | — | 71.4 | 72.6 | — | 79.6 | — | — | — | high |
81— | GPT-5.5 OpenAI🇺🇸5mo | 70.9 | 52.3 | 51.0 | 89.7 | 80.8 | — | 86.6 | $30 | — | — | high |
82— | GPT-5 Codex OpenAI🇺🇸1y | 70.7 | 70.7 | — | — | — | — | — | — | — | — | low |
83— | Claude Opus 4.1 Anthropic🇺🇸1y | 70 | 75.0 | 47.6 | 74.4 | 74.3 | — | — | — | — | — | medium |
84— | S Step3-VL-10B StepFun🇨🇳8moOSS | 70.0 | — | 65.9 | — | — | — | 76.1 | — | — | — | low |
85— | OpenAI Codex OpenAI🇺🇸1y | 69.7 | 58.4 | — | — | — | — | — | — | — | — | low |
86— | GPT-5.5 Instant OpenAI🇺🇸4mo | 69.7 | — | 53.8 | 81.4 | — | — | 70.1 | $30 | — | — | medium |
87— | Q Qwen3.8-27B Alibaba Cloud / Qwen Team🇨🇳1moOSS | 69.5 | 64.3 | — | 66.5 | — | — | 86 | $3 | — | — | medium |
88— | B Seed 1.8 ByteDance🇨🇳7mo | 69.2 | 67.2 | 76.8 | — | — | — | 63.7 | $2 | — | — | medium |
89— | Q Qwen3.5-122B-A10B Alibaba Cloud / Qwen Team🇨🇳7moOSS | 69.0 | 65.2 | — | 81.5 | 64.0 | — | 70.3 | $2 | — | — | high |
90— | Gemini 3 Flash Google🇺🇸9mo | 69.0 | 77.6 | 84.0 | 65.0 | 62.0 | — | 69.5 | $3 | — | — | high |
91— | GPT-5.1 Codex OpenAI🇺🇸10mo | 69.0 | 69.0 | — | — | — | — | — | — | — | — | low |
92— | MAI-Thinking-1 Microsoft🇺🇸4mo | 68.8 | 49.3 | 80.6 | 79.4 | — | — | — | — | — | — | medium |
93— | MiniMax M3 MiniMax🇨🇳4moOSS | 68.7 | 67.9 | — | — | — | — | 76.9 | $1 | — | — | medium |
94— | Claude Sonnet 4 Anthropic🇺🇸1y | 68.3 | 53.9 | 66.8 | 70.5 | 90.0 | — | 62.3 | $15 | — | — | high |
95— | Q Qwen3.5-27B Alibaba Cloud / Qwen Team🇨🇳7moOSS | 68.2 | 66.1 | — | 81.6 | 60.8 | — | 70 | $3 | — | — | high |
96— | U Solar Pro 4 Upstage🇰🇷1mo | 67.9 | 62.0 | — | 85.6 | — | — | — | $1 | — | — | low |
97— | M Muse Spark Meta🇺🇸5mo | 67.7 | 52.6 | — | 75.0 | — | — | 83.3 | — | — | — | high |
98— | X Grok 4 Fast xAI🇺🇸1y | 67.7 | — | 73.2 | 53.3 | 98.7 | — | — | — | — | — | medium |
99— | GPT OSS 120B High OpenAI🇺🇸1yOSS | 67.6 | — | 74.0 | 74.4 | 52.2 | — | — | — | — | — | low |
100— | GPT-5.4 OpenAI🇺🇸7mo | 67.6 | 48.7 | 69.1 | 80.0 | — | — | 83.2 | $15 | — | — | high |
Yes. Click any model to see its full benchmark profile, or use the comparison feature to compare up to 4 models side by side with radar charts and benchmark-by-benchmark scoring.