Compare AI model performance across 48+ benchmarks. Our composite index aggregates coding, math, reasoning, and language scores into a single intelligence ranking.
At-a-glance rankings across the three dimensions that matter most.
Neura Intelligence Index; Higher is better
Max input tokens; Higher is better
USD per 1M output tokens; Lower is better
Ranked by the Neura Intelligence Index — a weighted composite of 48 benchmarks across 8 categories.
Find the sweet spot — models in the top-left quadrant offer the best value.
Top performers in each benchmark category.
Who leads on each individual benchmark — click any card to see full results.
Different tasks need different strengths. These indices re-weight our benchmarks for specific workflows.
Best for software development, code generation, and debugging
Best for scientific research, data analysis, and complex reasoning
Best for writing, editing, summarization, and creative tasks
Highest intelligence per dollar — the cost-efficiency sweet spot
The Neura Intelligence Index is a composite score (0-100) computed from 48+ individual benchmarks spanning 8 categories: coding, math, reasoning, general knowledge, language, multimodal, safety, and agentic tasks.
Not all models have scores on all benchmarks. The confidence indicator reflects benchmark coverage: high (>70% of benchmarks), medium (40-70%), or low (<40%). Weights are renormalized across available categories so models aren't penalized for missing data.
Scores are aggregated from official model cards, Papers With Code, HuggingFace Open LLM Leaderboard, LiveBench, and LMSYS Chatbot Arena. Each score includes a verification status (official, self-reported, or aggregated).
The Neura Intelligence Index is a composite score (0-100) that aggregates AI model performance across 15+ benchmarks spanning coding, math, reasoning, general knowledge, language, multimodal, safety, and agentic tasks. It uses min-max normalization and weighted category averaging to produce a single comparable score.
Benchmark scores are synced daily via an automated pipeline that aggregates data from official model cards, Papers With Code, HuggingFace Open LLM Leaderboard, LiveBench, and LMSYS Chatbot Arena.
Confidence reflects how many benchmarks a model has been tested on relative to the total. High means >70% benchmark coverage, medium is 40-70%, and low is <40%. Models with low coverage may have composite scores that shift as more benchmarks are added.
Category weights reflect real-world demand: Coding and Reasoning each get 20%, General and Math each get 15%, Language and Multimodal each get 10%, and Safety and Agent each get 5%. Weights are renormalized across available categories so models are not penalized for missing data.
| # | Model | Coverage | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
2— | Claude Fable 5 Anthropic🇺🇸1mo | 97.7 | 98.1 | — | 97.2 | — | — | — | $50 | — | — | low |
3— | Claude Mythos Preview Anthropic🇺🇸 | 92.1 | 97.7 | — | 94.6 | 83.3 | — | 91.8 | — | — | — | medium |
4— | Claude Opus 4.8 Anthropic🇺🇸1mo | 91.9 | 93.2 | — | 92.5 | — | — | 91.9 | $25 | — | — | high |
5— | GPT-5.6 Sol OpenAI🇺🇸1w | 90.8 | 82.5 | 98.2 | 91.9 | 97.5 | — | 86.6 | $30 | — | — | medium |
6— | M Kimi K3 Moonshot AI🇨🇳5dOSS | 90.6 | — | — | 91.8 | — | — | 87.0 | $15 | — | — | high |
9— | GPT-5.6 Terra OpenAI🇺🇸1w | 88.1 | 77.6 | 97.4 | 90.5 | 97.0 | — | 82.6 | $15 | — | — | medium |
10— | X Grok 4.5 xAI🇺🇸5d | 86.7 | 82.8 | — | 90.6 | — | — | — | $6 | — | — | low |
11— | Claude Sonnet 5 Anthropic🇺🇸3w | 86.6 | 82.5 | — | 93.5 | — | — | 85.6 | $15 | — | — | medium |
12— | Claude Opus 4.7 Anthropic🇺🇸3mo | 86.0 | 86.1 | — | 91.5 | 80.1 | — | 89.3 | $25 | — | — | high |
13— | X Grok-4 Heavy xAI🇺🇸 | 85.6 | — | 84.0 | 86.7 | — | — | — | — | — | — | low |
14— | GPT-5.1 High OpenAI🇺🇸8mo | 84.9 | — | 83.5 | 85.8 | — | — | — | — | — | — | low |
15— | M Muse Spark 1.1 Meta🇺🇸1w | 84.3 | 68.5 | — | 96.3 | — | — | 85.7 | $4 | — | — | medium |
16— | GPT-5.1 Medium OpenAI🇺🇸8mo | 82.1 | — | 82.1 | — | — | — | — | $10 | — | — | low |
17— | B Seed 2.1 Pro ByteDance🇨🇳3w | 82 | 71.0 | — | 92.2 | — | — | 84.3 | — | — | — | high |
18— | Claude Opus 4.6 Anthropic🇺🇸5mo | 81.6 | 82.4 | 83.8 | 87 | 84.8 | — | 69.1 | $25 | — | — | high |
19— | GPT-5 High OpenAI🇺🇸11mo | 81.5 | — | 77.0 | 84.9 | — | — | — | — | — | — | low |
21— | ChatGPT-4o Latest OpenAI🇺🇸2y | 80.7 | — | — | 80.7 | — | — | — | — | — | — | low |
22— | B Seed 2.0 Pro ByteDance🇨🇳5mo | 80.5 | 74.9 | 82.0 | 86.7 | — | — | — | $3 | — | — | medium |
23— | GPT-5.1 Codex High OpenAI🇺🇸8mo | 79.9 | — | 79.9 | — | — | — | — | — | — | — | low |
24— | GPT-5.5 Pro OpenAI🇺🇸2mo | 79.5 | — | 57.5 | 93.4 | — | — | — | — | — | — | low |
25— | Claude Sonnet 4.5 Anthropic🇺🇸9mo | 79.1 | 92.8 | 64.7 | 79.9 | 72.5 | — | — | $15 | — | — | medium |
26— | B Seed 2.1 Turbo ByteDance🇨🇳3w | 79.0 | 68.3 | — | 91.3 | — | — | 81.2 | — | — | — | high |
28— | Z GLM-5.2 Zhipu AI🇨🇳1moOSS | 78.5 | 71.6 | — | 90.2 | — | — | — | $3 | — | — | medium |
29— | X Grok-3 xAI🇺🇸1y | 78.4 | — | 75.0 | 81.5 | — | — | 77.0 | $15 | — | — | low |
30— | GPT-5 Medium OpenAI🇺🇸11mo | 78.2 | — | 68.0 | 85.8 | — | — | — | — | — | — | low |
31— | X MiMo-V2-Pro Xiaomi🇨🇳4mo | 77.7 | 77.7 | — | — | — | — | — | — | — | — | low |
32— | Q Qwen3.7 Max Alibaba Cloud / Qwen Team🇨🇳2mo | 77.6 | 76.5 | — | 81.6 | 76.5 | — | — | $4 | — | — | medium |
33— | Z GLM-5V-Turbo Zhipu AI🇨🇳3mo | 77.3 | — | — | — | — | — | — | — | — | — | low |
34— | GPT-5.6 Luna OpenAI🇺🇸1w | 77.3 | 74.5 | 95.5 | 90.0 | 46.2 | — | 77.9 | $6 | — | — | medium |
35— | ERNIE 5.0 Baidu🇨🇳6mo | 77.2 | — | 64.7 | 75.3 | 92.2 | — | — | — | — | — | medium |
36— | GPT-5.2 Pro OpenAI🇺🇸7mo | 77.1 | — | 84.0 | 72.7 | — | — | — | — | — | — | medium |
37— | X Grok-3 Mini xAI🇺🇸1y | 76.6 | — | 71.1 | 80.7 | — | — | — | — | — | — | low |
38— | GPT-5.1 OpenAI🇺🇸8mo | 76.3 | 74.5 | 57.0 | 85.8 | — | — | 89.6 | $10 | — | — | medium |
39— | GPT-5.1 Instant OpenAI🇺🇸8mo | 76.3 | 74.5 | 57.0 | 85.8 | — | — | 89.6 | $10 | — | — | medium |
40— | GPT-5.1 Thinking OpenAI🇺🇸8mo | 76.3 | 74.5 | 57.0 | 85.8 | — | — | 89.6 | — | — | — | medium |
41— | Claude Sonnet 4.6 Anthropic🇺🇸5mo | 76.2 | 80.4 | — | 80.8 | 73.2 | — | 71.3 | $15 | — | — | high |
42— | o1 OpenAI🇺🇸1y | 75.7 | 49.0 | 94.2 | 80.3 | 86.7 | — | — | $60 | — | — | high |
43— | Claude Opus 4 Anthropic🇺🇸1y | 75.4 | 58.6 | 75.5 | 80.3 | 96.0 | — | 68.1 | $75 | — | — | high |
44— | B Seed 2.0 Lite ByteDance🇨🇳5mo | 75.3 | 68.9 | 74.6 | 82.2 | — | — | — | — | — | — | low |
45— | M Kimi K2-Thinking-0905 Moonshot AI🇨🇳10moOSS | 74.8 | 66.6 | 84.0 | 84.5 | — | — | — | — | — | — | medium |
46— | M Kimi K2.6 Moonshot AI🇨🇳3moOSS | 73.4 | 71.2 | — | 75.8 | — | — | 82.2 | $4 | — | — | high |
47— | T Hy3 Tencent🇨🇳2wOSS | 73.3 | 63.0 | — | 88.3 | — | — | — | — | — | — | medium |
48— | GPT-5.2 OpenAI🇺🇸7mo | 73.3 | 81.1 | 71.3 | 70.3 | 74.2 | — | 83.1 | $14 | — | — | high |
49— | Gemini 3.1 Pro Google🇺🇸5mo | 73.3 | 68.4 | — | 89.6 | 53.5 | — | 82.2 | $15 | — | — | high |
50— | GPT-5.5 OpenAI🇺🇸2mo | 73.2 | 52.4 | 51.0 | 91.3 | 89.3 | — | 86.9 | $30 | — | — | high |
51— | DeepSeek-V4-Pro-Max DeepSeek🇨🇳2moOSS | 73.2 | 58.2 | — | 86.0 | 77.7 | — | — | $3 | — | — | high |
52— | Z GLM-5 Zhipu AI🇨🇳5moOSS | 73.2 | 77.3 | — | — | — | — | — | $3 | — | — | low |
53— | o1-pro OpenAI🇺🇸1y | 73.1 | — | — | 73.1 | — | — | — | — | — | — | low |
55— | S Step-3.5-Flash StepFun🇨🇳5moOSS | 73.0 | 70.8 | 80.7 | — | — | — | — | $0.40 | — | — | low |
56— | Gemini 3 Pro Google🇺🇸8mo | 72.8 | 74.3 | 84.0 | 66.5 | 65.1 | — | 77.2 | — | — | — | high |
57— | GPT OSS 20B High OpenAI🇺🇸11moOSS | 72.3 | — | 82.5 | 64.7 | — | — | — | — | — | — | low |
58— | Claude Opus 4.5 Anthropic🇺🇸7mo | 72.0 | 82.5 | — | 61.5 | 78.0 | — | — | — | — | — | medium |
59— | Q Qwen3.7-Plus Alibaba Cloud / Qwen Team🇨🇳1mo | 71.7 | 66.8 | — | 73.8 | 72.2 | — | 82.7 | $1 | — | — | high |
60— | X MiMo-V2-Omni Xiaomi🇨🇳4mo | 71.6 | 71.6 | — | — | — | — | — | — | — | — | low |
61— | M Kimi K2.5 Moonshot AI🇨🇳5moOSS | 71.3 | 51.8 | 79.1 | 86.0 | — | — | 70.6 | — | — | — | high |
62— | GPT-5 Codex OpenAI🇺🇸10mo | 71.0 | 71.0 | — | — | — | — | — | — | — | — | low |
63— | GPT-5.5 Instant OpenAI🇺🇸2mo | 70.9 | — | 54.0 | 82.8 | — | — | 72.6 | $30 | — | — | medium |
64— | Claude Opus 4.1 Anthropic🇺🇸11mo | 70.5 | 75.2 | 47.9 | 76.2 | 73.8 | — | — | — | — | — | medium |
65— | S Step3-VL-10B StepFun🇨🇳6moOSS | 70.4 | — | 65.9 | — | — | — | 77.2 | — | — | — | low |
66— | Q Qwen3.5-122B-A10B Alibaba Cloud / Qwen Team🇨🇳4moOSS | 70.3 | 65.6 | — | 83.6 | 63.7 | — | 73.8 | — | — | — | high |
67— | Gemini 3 Flash Google🇺🇸7mo | 70.3 | 77.7 | 83.7 | 66.0 | 62.6 | — | 73.7 | $3 | — | — | high |
69— | Q Qwen3.5-27B Alibaba Cloud / Qwen Team🇨🇳4moOSS | 69.4 | 66.5 | — | 83.6 | 60.6 | — | 73.4 | $2 | — | — | high |
70— | GPT-5.1 Codex OpenAI🇺🇸8mo | 69.3 | 69.3 | — | — | — | — | — | — | — | — | low |
71— | Gemma 4 31B Google🇺🇸3moOSS | 69.2 | — | — | 61.2 | 76.3 | — | 74.5 | $0.38 | — | — | medium |
72— | MiniMax M3 MiniMax🇨🇳1moOSS | 69.1 | 68.3 | — | — | — | — | 77.2 | $1 | — | — | medium |
73— | Claude Sonnet 4 Anthropic🇺🇸1y | 68.9 | 54.5 | 66.8 | 71.6 | 90.0 | — | 63.7 | $15 | — | — | high |
74— | X Grok 4 Fast xAI🇺🇸10mo | 68.5 | — | 73.0 | 55.4 | 98.6 | — | — | — | — | — | medium |
75— | GPT-5.4 OpenAI🇺🇸4mo | 68.3 | 47.2 | 69.1 | 82.2 | — | — | 83.5 | $15 | — | — | high |
76— | GPT OSS 120B High OpenAI🇺🇸11moOSS | 68.3 | — | 73.8 | 76.2 | 52.1 | — | — | — | — | — | low |
77— | M Kimi K2.7 Code Moonshot AI🇨🇳1moOSS | 68.1 | — | — | — | — | — | — | $4 | — | — | low |
78— | DeepSeek-V3.2 DeepSeek🇨🇳7moOSS | 67.9 | 68.0 | 74.8 | 75.2 | — | — | — | — | — | — | medium |
79— | MAI-Thinking-1 Microsoft🇺🇸1mo | 67.8 | 45.3 | 80.3 | 81.0 | — | — | — | — | — | — | medium |
80— | Q Qwen3.5-397B-A17B Alibaba Cloud / Qwen Team🇨🇳5moOSS | 67.7 | 74.7 | — | 66.1 | 70.4 | — | — | — | — | — | medium |
81— | M Kimi K2 0905 Moonshot AI🇨🇳10mo | 67.6 | — | — | 67.6 | — | — | — | — | — | — | low |
82— | Z GLM-4.7 Zhipu AI🇨🇳7moOSS | 67.3 | 57.0 | 78.5 | 79.3 | — | — | — | — | — | — | medium |
83— | L K-EXAONE-236B-A23B LG AI Research🇰🇷6mo | 67.0 | — | 74.3 | — | 59.8 | — | — | — | — | — | low |
84— | Nova 2 Pro Amazon🇺🇸7mo | 67.0 | 67.3 | 73.5 | 76.9 | — | — | 37.0 | — | — | — | medium |
85— | Z GLM-5.1 Zhipu AI🇨🇳3moOSS | 66.9 | 51.3 | — | 86.3 | — | — | — | $4 | — | — | medium |
86— | M Muse Spark Meta🇺🇸3mo | 66.8 | 48.3 | — | 75.8 | — | — | 85.9 | — | — | — | high |
87— | Gemini 2.0 Flash Thinking Google🇺🇸1y | 66.8 | — | — | 64.7 | — | — | 71.1 | — | — | — | low |
88— | Q Qwen3-Next-80B-A3B-Thinking Alibaba Cloud / Qwen Team🇨🇳10moOSS | 66.8 | — | 66.1 | 70.1 | — | — | — | — | — | — | low |
89— | Llama 3.1 405B Meta🇺🇸1yOSS | 66.3 | 69.3 | 72.0 | 48.9 | 67.9 | 87.7 | — | $3 | — | — | high |
90— | Q Qwen3.6 Plus Alibaba Cloud / Qwen Team🇨🇳3mo | 66.1 | 60.0 | — | 67.3 | 73.8 | — | 76.8 | $3 | — | — | high |
91— | Q Qwen3.5-35B-A3B Alibaba Cloud / Qwen Team🇨🇳4moOSS | 65.9 | 59.3 | — | 82.0 | 57.8 | — | 70.8 | — | — | — | high |
92— | Mistral Medium 3.5 Mistral AI🇫🇷2moOSS | 64.9 | 77.0 | 63.4 | — | — | — | — | $8 | — | — | low |
94— | Q Qwen3-Coder 480B A35B Instruct Alibaba Cloud / Qwen Team🇨🇳1yOSS | 63.6 | 60.2 | — | — | — | — | — | — | — | — | low |
95— | GPT-5 OpenAI🇺🇸11mo | 63.5 | 71.8 | 57.1 | 60.3 | — | — | 79.2 | — | — | — | high |
96— | DeepSeek-V3.2-Exp DeepSeek🇨🇳9moOSS | 63.5 | 58.6 | 68.6 | 51.0 | 98.8 | — | — | — | — | — | medium |
97— | X Grok-4 xAI🇺🇸1y | 63.2 | — | 72.6 | 56.2 | — | — | — | — | — | — | medium |
98— | M LongCat-Flash-Thinking-2601 Meituan🇨🇳6moOSS | 63.1 | 61.1 | 83.5 | 57.0 | — | — | — | — | — | — | medium |
99— | DeepSeek R1 Zero DeepSeek🇨🇳1yOSS | 63.0 | — | — | 63.0 | — | — | — | — | — | — | low |
100— | X Grok Code Fast 1 xAI🇺🇸10mo | 63.0 | 63.0 | — | — | — | — | — | $2 | — | — | low |
101— | DeepSeek-V3.2 (Thinking) DeepSeek🇨🇳7moOSS | 62.3 | 68.0 | 74.8 | 58.3 | — | — | — | — | — | — | medium |
102— | Q Qwen3.5-9B Alibaba Cloud / Qwen Team🇨🇳4moOSS | 62.0 | — | — | 77.4 | 41.5 | — | — | — | — | — | low |
103— | MiniMax M2.5 MiniMax🇨🇳5moOSS | 60.6 | 57.8 | — | — | — | — | — | $1 | — | — | low |
105— | M LongCat-Flash-Thinking Meituan🇨🇳10moOSS | 60.5 | 36.3 | 70.8 | 77.1 | — | — | — | — | — | — | low |
106— | Claude 3.5 Sonnet Anthropic🇺🇸2y | 60.4 | 48.4 | 58.6 | 68.4 | 80.5 | 31.6 | 52.2 | $15 | — | — | high |
107— | Q Qwen3-235B-A22B-Instruct-2507 Alibaba Cloud / Qwen Team🇨🇳12moOSS | 60.4 | — | 33.7 | 70.6 | 73.3 | — | — | — | — | — | low |
108— | Gemini 3.1 Flash-Lite Google🇺🇸4mo | 60.1 | — | — | 52.6 | 68.2 | — | 63.0 | $2 | — | — | medium |
109— | DeepSeek-V3.2-Speciale DeepSeek🇨🇳7moOSS | 60.0 | 68.0 | 79.0 | 50.4 | — | — | — | — | — | — | medium |
Yes. Click any model to see its full benchmark profile, or use the comparison feature to compare up to 4 models side by side with radar charts and benchmark-by-benchmark scoring.