Compare AI model performance across 48+ benchmarks. Our composite index aggregates coding, math, reasoning, and language scores into a single intelligence ranking.
At-a-glance rankings across the three dimensions that matter most.
Neura Intelligence Index; Higher is better
Max input tokens; Higher is better
USD per 1M output tokens; Lower is better
Ranked by the Neura Intelligence Index — a weighted composite of 48 benchmarks across 8 categories.
Find the sweet spot — models in the top-left quadrant offer the best value.
Top performers in each benchmark category.
Who leads on each individual benchmark — click any card to see full results.
Different tasks need different strengths. These indices re-weight our benchmarks for specific workflows.
Best for software development, code generation, and debugging
Best for scientific research, data analysis, and complex reasoning
Best for writing, editing, summarization, and creative tasks
Highest intelligence per dollar — the cost-efficiency sweet spot
The Neura Intelligence Index is a composite score (0-100) computed from 48+ individual benchmarks spanning 8 categories: coding, math, reasoning, general knowledge, language, multimodal, safety, and agentic tasks.
Not all models have scores on all benchmarks. The confidence indicator reflects benchmark coverage: high (>70% of benchmarks), medium (40-70%), or low (<40%). Weights are renormalized across available categories so models aren't penalized for missing data.
Scores are aggregated from official model cards, Papers With Code, HuggingFace Open LLM Leaderboard, LiveBench, and LMSYS Chatbot Arena. Each score includes a verification status (official, self-reported, or aggregated).
The Neura Intelligence Index is a composite score (0-100) that aggregates AI model performance across 15+ benchmarks spanning coding, math, reasoning, general knowledge, language, multimodal, safety, and agentic tasks. It uses min-max normalization and weighted category averaging to produce a single comparable score.
Benchmark scores are synced daily via an automated pipeline that aggregates data from official model cards, Papers With Code, HuggingFace Open LLM Leaderboard, LiveBench, and LMSYS Chatbot Arena.
Confidence reflects how many benchmarks a model has been tested on relative to the total. High means >70% benchmark coverage, medium is 40-70%, and low is <40%. Models with low coverage may have composite scores that shift as more benchmarks are added.
Category weights reflect real-world demand: Coding and Reasoning each get 20%, General and Math each get 15%, Language and Multimodal each get 10%, and Safety and Agent each get 5%. Weights are renormalized across available categories so models are not penalized for missing data.
| # | Model | Coverage | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
1— | X Grok 4.6 xAI🇺🇸3w | 99.4 | — | — | — | — | — | — | $6 | — | — | low |
3— | Claude Fable 5 Anthropic🇺🇸2mo | 97.2 | 98.0 | — | 96.3 | — | — | — | $50 | — | — | low |
4— | Claude Fable 5.1 Anthropic🇺🇸4d | 96.6 | — | — | 96.6 | — | — | — | $50 | — | — | low |
5— | M Muse Spark 1.3 Meta🇺🇸3d | 95.7 | — | — | — | 95.7 | — | — | $0.20 | — | — | low |
6— | Claude Opus 5 Anthropic🇺🇸1mo | 95.1 | — | — | 96.4 | — | — | — | $25 | — | — | low |
7— | Z GLM-5.3 Zhipu AI🇨🇳3wOSS | 94.8 | — | — | 95.4 | — | — | — | $4 | — | — | low |
8— | GPT-6 Astra OpenAI🇺🇸1d | 94.2 | — | — | 94.2 | — | — | 96.3 | $50 | — | — | medium |
9— | DeepSeek-V4-Pro-0813 DeepSeek🇨🇳3wOSS | 93.8 | — | — | 93.9 | — | — | — | $0.87 | — | — | low |
10— | Claude Mythos Preview Anthropic🇺🇸 | 91.4 | 97.4 | — | 93.8 | 83.3 | — | 89.5 | — | — | — | medium |
11— | Gemini 3.7 Flash Google🇺🇸3w | 90.2 | — | — | — | 95.2 | — | 82.8 | $4 | — | — | low |
12— | Claude Opus 4.8 Anthropic🇺🇸3mo | 89.8 | 91.7 | — | 91.3 | — | — | 88.9 | $25 | — | — | high |
13— | DeepSeek-V4-Flash-0731 DeepSeek🇨🇳1moOSS | 89.4 | — | — | — | — | — | — | $0.18 | — | — | low |
15— | Z GLM-5.3-Flash Zhipu AI🇨🇳1wOSS | 89.3 | — | — | 90.0 | — | — | 84.0 | $0.50 | — | — | low |
16— | GPT-5.6 Sol OpenAI🇺🇸1mo | 88.7 | 80.0 | 98.2 | 91.2 | 93.0 | — | 86.2 | $30 | — | — | medium |
18— | M Kimi K3 Moonshot AI🇨🇳1moOSS | 88.0 | — | — | 90.5 | — | — | 85.4 | $15 | — | — | high |
19— | Q Qwen3.8 Max Alibaba Cloud / Qwen Team🇨🇳1moOSS | 87.7 | 88.3 | — | 81.7 | 93.6 | — | 87.3 | $5 | — | — | medium |
20— | GPT-5.6 Terra OpenAI🇺🇸1mo | 86.1 | 76.1 | 97.4 | 89.8 | 92.0 | — | 82.2 | $12 | — | — | medium |
21— | T Hy4 preview Tencent🇨🇳1wOSS | 85.6 | 83.3 | — | 89.2 | — | — | — | — | — | — | medium |
22— | X Grok 4.5 xAI🇺🇸1mo | 85.1 | 80.3 | — | 89.8 | — | — | — | $6 | — | — | low |
23— | Claude Opus 4.7 Anthropic🇺🇸4mo | 84.8 | 85.1 | — | 90.2 | 80.1 | — | 86.5 | $25 | — | — | high |
24— | X Grok-4 Heavy xAI🇺🇸 | 84.7 | — | 84.2 | 85.0 | — | — | — | — | — | — | low |
25— | Claude Sonnet 5 Anthropic🇺🇸2mo | 84.3 | 82.0 | — | 91.9 | — | — | 82.0 | $10 | — | — | medium |
26— | GPT-5.1 High OpenAI🇺🇸9mo | 84.3 | — | 83.8 | 84.8 | — | — | — | — | — | — | low |
27— | M Muse Spark 1.1 Meta🇺🇸1mo | 83.2 | 69.0 | — | 95.2 | — | — | 82.2 | $4 | — | — | medium |
28— | GPT-5.1 Medium OpenAI🇺🇸9mo | 82.3 | — | 82.3 | — | — | — | — | $10 | — | — | low |
29— | B Seed 2.1 Pro ByteDance🇨🇳2mo | 81.3 | 74.5 | — | 90.4 | — | — | 82.0 | — | — | — | high |
30— | GPT-5 High OpenAI🇺🇸1y | 81.0 | — | 77.2 | 83.8 | — | — | — | — | — | — | low |
32— | B Seed 2.0 Pro ByteDance🇨🇳6mo | 80.2 | 75.1 | 82.2 | 85.7 | — | — | — | $3 | — | — | medium |
33— | GPT-5.1 Codex High OpenAI🇺🇸9mo | 80.1 | — | 80.1 | — | — | — | — | — | — | — | low |
34— | Claude Opus 4.6 Anthropic🇺🇸7mo | 80.0 | 82.6 | 84 | 84.8 | 80.6 | — | 65.9 | $25 | — | — | high |
35— | ChatGPT-4o Latest OpenAI🇺🇸2y | 79.4 | — | — | 79.4 | — | — | — | — | — | — | low |
36— | B Seed 2.1 Turbo ByteDance🇨🇳2mo | 78.9 | 72.4 | — | 89.4 | — | — | 78.6 | — | — | — | high |
38— | Claude Sonnet 4.5 Anthropic🇺🇸11mo | 78.7 | 92.8 | 64.8 | 78.5 | 72.5 | — | — | $15 | — | — | medium |
39— | GPT-5.5 Pro OpenAI🇺🇸4mo | 78.6 | — | 57.5 | 91.8 | — | — | — | — | — | — | low |
40— | Gemini 3.8 Flash Google🇺🇸3d | 78.0 | — | — | — | — | — | 78.0 | $4 | — | — | low |
41— | X MiMo-V2-Pro Xiaomi🇨🇳5mo | 77.9 | 77.9 | — | — | — | — | — | — | — | — | low |
42— | X Grok-3 xAI🇺🇸1y | 77.8 | — | 75.2 | 80.3 | — | — | 76.7 | $15 | — | — | low |
43— | GPT-5 Medium OpenAI🇺🇸1y | 77.6 | — | 68.1 | 84.8 | — | — | — | — | — | — | low |
44— | Q Qwen3.7 Max Alibaba Cloud / Qwen Team🇨🇳3mo | 77.4 | 78.3 | — | 79.5 | 76.5 | — | — | $4 | — | — | medium |
45— | Z GLM-5V-Turbo Zhipu AI🇨🇳5mo | 77.3 | — | — | — | — | — | — | — | — | — | low |
46— | Q Qwen3.8 Flash Alibaba Cloud / Qwen Team🇨🇳1w | 77.3 | 72.9 | — | 73.5 | — | — | 85.9 | $0.47 | — | — | medium |
47— | Q Qwen3.8-Flash-Next Alibaba Cloud / Qwen Team🇨🇳1wOSS | 77.3 | 72.9 | — | 73.5 | — | — | 85.9 | — | — | — | medium |
48— | Z GLM-5.2 Zhipu AI🇨🇳2moOSS | 76.9 | 71.4 | — | 88.8 | — | — | — | $3 | — | — | medium |
49— | ERNIE 5.0 Baidu🇨🇳7mo | 76.3 | — | 64.8 | 72.9 | 92.5 | — | — | — | — | — | medium |
50— | GPT-5.1 OpenAI🇺🇸9mo | 76.0 | 74.7 | 57.0 | 84.8 | — | — | 89.5 | $10 | — | — | medium |
51— | GPT-5.1 Instant OpenAI🇺🇸9mo | 76.0 | 74.7 | 57.0 | 84.8 | — | — | 89.5 | $10 | — | — | medium |
52— | GPT-5.1 Thinking OpenAI🇺🇸9mo | 76.0 | 74.7 | 57.0 | 84.8 | — | — | 89.5 | — | — | — | medium |
53— | X Grok-3 Mini xAI🇺🇸1y | 75.9 | — | 71.3 | 79.4 | — | — | — | — | — | — | low |
54— | GPT-5.2 Pro OpenAI🇺🇸8mo | 75.8 | — | 84.2 | 69.9 | — | — | — | — | — | — | medium |
55— | Claude Sonnet 4.6 Anthropic🇺🇸6mo | 75.5 | 80.7 | — | 78.2 | 73.2 | — | 71.1 | $15 | — | — | high |
56— | o1 OpenAI🇺🇸1y | 75.3 | 48.6 | 94.2 | 79.5 | 86.7 | — | — | $60 | — | — | high |
57— | M Kimi K2-Thinking-0905 Moonshot AI🇨🇳12moOSS | 75.3 | 69.6 | 84.2 | 82.7 | — | — | — | — | — | — | medium |
58— | Claude Opus 4 Anthropic🇺🇸1y | 75.1 | 58.3 | 75.5 | 79.5 | 96.0 | — | 67.8 | $75 | — | — | high |
59— | GPT-5.6 Luna OpenAI🇺🇸1mo | 74.9 | 73.6 | 95.5 | 89.2 | 38.4 | — | 77.6 | $1 | — | — | medium |
60— | B Seed 2.0 Lite ByteDance🇨🇳6mo | 74.9 | 68.9 | 74.8 | 81.0 | — | — | — | — | — | — | low |
61— | T Hy3 Tencent🇨🇳2moOSS | 73.9 | 65.8 | — | 87.3 | — | — | — | — | — | — | medium |
62— | DeepSeek-V4-Pro-Max DeepSeek🇨🇳4moOSS | 73.6 | 62.5 | — | 84.2 | 78.1 | — | — | — | — | — | high |
63— | S Step-3.5-Flash StepFun🇨🇳7moOSS | 73.0 | 70.8 | 80.9 | — | — | — | — | $0.40 | — | — | low |
65— | Z GLM-5 Zhipu AI🇨🇳6moOSS | 73 | 77.5 | — | — | — | — | — | $3 | — | — | low |
66— | M Kimi K2.6 Moonshot AI🇨🇳4moOSS | 72.8 | 74.3 | — | 73.5 | — | — | 80.0 | $4 | — | — | high |
67— | Gemini 3.1 Pro Google🇺🇸6mo | 72.6 | 72.0 | — | 87.7 | 51.8 | — | 81.8 | $15 | — | — | high |
68— | L EXAONE 4.5 33B LG AI Research🇰🇷4mo | 72.3 | — | 74.6 | 74.1 | — | — | 65.2 | — | — | — | medium |
69— | GPT-5.2 OpenAI🇺🇸8mo | 72.1 | 81.3 | 71.4 | 67.5 | 74.2 | — | 80.1 | $14 | — | — | high |
70— | M Kimi K2.5 Moonshot AI🇨🇳7moOSS | 72.0 | 57.5 | 79.3 | 84.2 | — | — | 67.3 | — | — | — | high |
71— | GPT-5.5 OpenAI🇺🇸4mo | 71.8 | 56.9 | 51.0 | 89.6 | 80.4 | — | 86.5 | $30 | — | — | high |
72— | Gemini 3 Pro Google🇺🇸9mo | 71.7 | 74.5 | 84.2 | 64.5 | 64.1 | — | 73 | — | — | — | high |
73— | X MiMo-V2-Omni Xiaomi🇨🇳5mo | 71.7 | 71.7 | — | — | — | — | — | — | — | — | low |
74— | Q Qwen3.8-27B Alibaba Cloud / Qwen Team🇨🇳3wOSS | 71.7 | 69.8 | — | 66.7 | — | — | 85.3 | — | — | — | medium |
75— | o1-pro OpenAI🇺🇸1y | 71.6 | — | — | 71.6 | — | — | — | — | — | — | low |
76— | Q Qwen3.7-Plus Alibaba Cloud / Qwen Team🇨🇳3mo | 71.5 | 70.5 | — | 71.5 | 72.2 | — | 79.3 | — | — | — | high |
77— | GPT OSS 20B High OpenAI🇺🇸1yOSS | 71.4 | — | 82.7 | 62.9 | — | — | — | — | — | — | low |
78— | Claude Opus 4.5 Anthropic🇺🇸9mo | 71.3 | 82.8 | — | 59.5 | 78.0 | — | — | — | — | — | medium |
79— | GPT-5 Codex OpenAI🇺🇸11mo | 71.0 | 71.0 | — | — | — | — | — | — | — | — | low |
80— | S Step3-VL-10B StepFun🇨🇳7moOSS | 70.4 | — | 66 | — | — | — | 76.9 | — | — | — | low |
81— | Claude Opus 4.1 Anthropic🇺🇸1y | 70.1 | 75.2 | 47.8 | 74.7 | 73.8 | — | — | — | — | — | medium |
82— | MiniMax M3 MiniMax🇨🇳3moOSS | 70.1 | 70.4 | — | — | — | — | 76.9 | $1 | — | — | medium |
84— | GPT-5.5 Instant OpenAI🇺🇸4mo | 69.8 | — | 54.0 | 81.6 | — | — | 69.8 | $30 | — | — | medium |
85— | GPT-5.1 Codex OpenAI🇺🇸9mo | 69.3 | 69.3 | — | — | — | — | — | — | — | — | low |
86— | MAI-Thinking-1 Microsoft🇺🇸3mo | 69.3 | 50.4 | 80.5 | 79.7 | — | — | — | — | — | — | medium |
87— | Q Qwen3.5-122B-A10B Alibaba Cloud / Qwen Team🇨🇳6moOSS | 69.2 | 65.6 | — | 81.6 | 63.7 | — | 70.4 | — | — | — | high |
88— | Gemini 3 Flash Google🇺🇸8mo | 68.7 | 77.9 | 83.9 | 63.7 | 62.0 | — | 69.3 | $3 | — | — | high |
89— | GPT-5.4 OpenAI🇺🇸6mo | 68.7 | 52.9 | 69.1 | 79.7 | — | — | 83.1 | $15 | — | — | high |
90— | Claude Sonnet 4 Anthropic🇺🇸1y | 68.5 | 54.2 | 66.8 | 70.7 | 90.0 | — | 63.3 | $15 | — | — | high |
91— | Q Qwen3.5-27B Alibaba Cloud / Qwen Team🇨🇳6moOSS | 68.4 | 66.5 | — | 81.7 | 60.6 | — | 70.1 | $2 | — | — | high |
92— | U Solar Pro 4 Upstage🇰🇷4w | 68.3 | 62.4 | — | 85.8 | — | — | — | $1 | — | — | low |
93— | Z GLM-5.1 Zhipu AI🇨🇳5moOSS | 67.9 | 56.0 | — | 84.6 | — | — | — | $4 | — | — | medium |
94— | X Grok 4 Fast xAI🇺🇸1y | 67.8 | — | 73.2 | 53.5 | 98.7 | — | — | — | — | — | medium |
95— | GPT OSS 120B High OpenAI🇺🇸1yOSS | 67.7 | — | 74 | 74.7 | 52.1 | — | — | — | — | — | low |
96— | M Muse Spark Meta🇺🇸5mo | 67.5 | 53.5 | — | 73.8 | — | — | 83.0 | — | — | — | high |
97— | L K-EXAONE-236B-A23B LG AI Research🇰🇷8mo | 67.1 | — | 74.5 | — | 59.8 | — | — | — | — | — | low |
98— | DeepSeek-V3.2 DeepSeek🇨🇳9moOSS | 67.1 | 68.0 | 74.9 | 72.8 | — | — | — | — | — | — | medium |
99— | MAI-Code-1.1-Flash Microsoft🇺🇸3w | 66.9 | 66.9 | — | — | — | — | — | $1 | — | — | low |
100— | Q Qwen3.5-397B-A17B Alibaba Cloud / Qwen Team🇨🇳6moOSS | 66.8 | 74.9 | — | 63.9 | 70.4 | — | — | — | — | — | medium |
101— | Nova 2 Pro Amazon🇺🇸9mo | 66.7 | 67.3 | 73.7 | 75.5 | — | — | 37.8 | — | — | — | medium |
102— | Z GLM-4.7 Zhipu AI🇨🇳8moOSS | 66.6 | 57.0 | 78.7 | 77.1 | — | — | — | — | — | — | medium |
103— | Gemma 4 31B Google🇺🇸5moOSS | 66.4 | — | — | 58.9 | 71.1 | — | 74.2 | $0.38 | — | — | medium |
104— | M Kimi K2.7 Code Moonshot AI🇨🇳2moOSS | 66.4 | — | — | — | — | — | — | — | — | — | low |
105— | Llama 3.1 405B Meta🇺🇸2yOSS | 66.2 | 69.3 | 72.0 | 48.4 | 67.9 | 87.7 | — | $3 | — | — | high |
106— | Q Qwen3-Next-80B-A3B-Thinking Alibaba Cloud / Qwen Team🇨🇳12moOSS | 66.0 | — | 66.2 | 68.4 | — | — | — | — | — | — | low |
107— | M Kimi K2 0905 Moonshot AI🇨🇳12mo | 65.9 | — | — | 65.9 | — | — | — | — | — | — | low |
Yes. Click any model to see its full benchmark profile, or use the comparison feature to compare up to 4 models side by side with radar charts and benchmark-by-benchmark scoring.