Compare AI model performance across 48+ benchmarks. Our composite index aggregates coding, math, reasoning, and language scores into a single intelligence ranking.
At-a-glance rankings across the three dimensions that matter most.
Neura Intelligence Index; Higher is better
Max input tokens; Higher is better
USD per 1M output tokens; Lower is better
Ranked by the Neura Intelligence Index β a weighted composite of 48 benchmarks across 8 categories.
Find the sweet spot β models in the top-left quadrant offer the best value.
Top performers in each benchmark category.
Who leads on each individual benchmark β click any card to see full results.
Different tasks need different strengths. These indices re-weight our benchmarks for specific workflows.
Best for software development, code generation, and debugging
Best for scientific research, data analysis, and complex reasoning
Best for writing, editing, summarization, and creative tasks
Highest intelligence per dollar β the cost-efficiency sweet spot
The Neura Intelligence Index is a composite score (0-100) computed from 48+ individual benchmarks spanning 8 categories: coding, math, reasoning, general knowledge, language, multimodal, safety, and agentic tasks.
Not all models have scores on all benchmarks. The confidence indicator reflects benchmark coverage: high (>70% of benchmarks), medium (40-70%), or low (<40%). Weights are renormalized across available categories so models aren't penalized for missing data.
Scores are aggregated from official model cards, Papers With Code, HuggingFace Open LLM Leaderboard, LiveBench, and LMSYS Chatbot Arena. Each score includes a verification status (official, self-reported, or aggregated).
The Neura Intelligence Index is a composite score (0-100) that aggregates AI model performance across 15+ benchmarks spanning coding, math, reasoning, general knowledge, language, multimodal, safety, and agentic tasks. It uses min-max normalization and weighted category averaging to produce a single comparable score.
Benchmark scores are synced daily via an automated pipeline that aggregates data from official model cards, Papers With Code, HuggingFace Open LLM Leaderboard, LiveBench, and LMSYS Chatbot Arena.
Confidence reflects how many benchmarks a model has been tested on relative to the total. High means >70% benchmark coverage, medium is 40-70%, and low is <40%. Models with low coverage may have composite scores that shift as more benchmarks are added.
Category weights reflect real-world demand: Coding and Reasoning each get 20%, General and Math each get 15%, Language and Multimodal each get 10%, and Safety and Agent each get 5%. Weights are renormalized across available categories so models are not penalized for missing data.
| # | Model | Coverage | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
1β | X Grok 4.6 xAIπΊπΈ2d | 99.3 | β | β | β | β | β | β | $6 | β | β | low |
2β | GPT-Image-1 OpenAIπΊπΈ1y | 98.7 | β | β | β | β | β | β | β | β | β | low |
3β | Claude Fable 5 AnthropicπΊπΈ2mo | 97.4 | 98.1 | β | 96.6 | β | β | β | $50 | β | β | low |
4β | Z GLM-5.3 Zhipu AIπ¨π³0dOSS | 95.8 | β | β | 95.7 | β | β | β | β | β | β | low |
5β | Claude Opus 5 AnthropicπΊπΈ3w | 95.5 | β | β | 96.7 | β | β | β | $25 | β | β | low |
6β | DeepSeek-V4-Pro-0813 DeepSeekπ¨π³1dOSS | 94.8 | β | β | 94.3 | β | β | β | $0.87 | β | β | low |
7β | DeepSeek-V4-Flash-0731 DeepSeekπ¨π³2wOSS | 94.1 | β | β | β | β | β | β | $0.18 | β | β | low |
8β | Gemini 3.7 Flash GoogleπΊπΈ1d | 92.1 | β | β | β | 96.6 | β | 85.3 | $4 | β | β | low |
9β | Claude Mythos Preview AnthropicπΊπΈ | 91.9 | 97.8 | β | 94.1 | 83.3 | β | 91.3 | β | β | β | medium |
10β | Claude Opus 4.8 AnthropicπΊπΈ2mo | 91.3 | 93.5 | β | 91.8 | β | β | 91.0 | $25 | β | β | high |
11β | GPT-5.6 Sol OpenAIπΊπΈ1mo | 90.2 | 83.7 | 98.2 | 91.5 | 94.8 | β | 86 | $30 | β | β | medium |
12β | Q Qwen3.8 Max Alibaba Cloud / Qwen Teamπ¨π³1wOSS | 89.9 | 92.5 | β | 82.5 | 95.3 | β | 88.4 | β | β | β | medium |
13β | Sora OpenAIπΊπΈ1y | 89.4 | β | β | β | β | β | β | β | β | β | low |
14β | M Kimi K3 Moonshot AIπ¨π³4wOSS | 88.8 | β | β | 91.0 | β | β | 86.4 | $15 | β | β | high |
15β | Veo 2 GoogleπΊπΈ1y | 88.2 | β | β | β | β | β | β | β | β | β | low |
16β | GPT-5.6 Terra OpenAIπΊπΈ1mo | 87.4 | 78.9 | 97.4 | 90.2 | 94.0 | β | 82.0 | $12 | β | β | medium |
17β | X Grok 4.5 xAIπΊπΈ4w | 87.2 | 84.1 | β | 90.2 | β | β | β | $6 | β | β | low |
18β | Claude Opus 4.7 AnthropicπΊπΈ4mo | 85.8 | 86.7 | β | 90.7 | 80.1 | β | 88.7 | $25 | β | β | high |
19β | Claude Sonnet 5 AnthropicπΊπΈ1mo | 85.8 | 83.2 | β | 92.5 | β | β | 84.7 | $10 | β | β | medium |
20β | X Grok-4 Heavy xAIπΊπΈ | 85.0 | β | 84.2 | 85.7 | β | β | β | β | β | β | low |
21β | GPT-5.1 High OpenAIπΊπΈ9mo | 84.6 | β | 83.8 | 85.3 | β | β | β | β | β | β | low |
22β | M Muse Spark 1.1 MetaπΊπΈ1mo | 84.3 | 69.9 | β | 95.5 | β | β | 84.8 | $4 | β | β | medium |
23β | GPT-5.1 Medium OpenAIπΊπΈ9mo | 82.3 | β | 82.3 | β | β | β | β | $10 | β | β | low |
24β | GPT-5 High OpenAIπΊπΈ1y | 81.3 | β | 77.3 | 84.4 | β | β | β | β | β | β | low |
25β | B Seed 2.1 Pro ByteDanceπ¨π³1mo | 81.0 | 71.9 | β | 91.0 | β | β | 83.4 | β | β | β | high |
26β | DALL-E 3 OpenAIπΊπΈ2y | 81.0 | β | β | β | β | β | β | β | β | β | low |
27β | Claude Opus 4.6 AnthropicπΊπΈ6mo | 80.9 | 82.2 | 84.0 | 86.7 | 82.1 | β | 67.8 | $25 | β | β | high |
28β | B Seed 2.0 Pro ByteDanceπ¨π³6mo | 80.3 | 74.6 | 82.2 | 86.2 | β | β | β | $3 | β | β | medium |
29β | GPT-5.1 Codex High OpenAIπΊπΈ9mo | 80.2 | β | 80.2 | β | β | β | β | β | β | β | low |
30β | ChatGPT-4o Latest OpenAIπΊπΈ2y | 80.1 | β | β | 80.1 | β | β | β | β | β | β | low |
31β | Claude Sonnet 4.5 AnthropicπΊπΈ10mo | 79.0 | 92.8 | 65.0 | 79.2 | 72.5 | β | β | $15 | β | β | medium |
32β | GPT-5.5 Pro OpenAIπΊπΈ3mo | 78.9 | β | 57.5 | 92.3 | β | β | β | β | β | β | low |
33β | Claude Code AnthropicπΊπΈ1y | 78.6 | 66.7 | β | β | β | β | β | β | β | β | low |
34β | B Seed 2.1 Turbo ByteDanceπ¨π³1mo | 78.5 | 69.3 | β | 90.0 | β | β | 80.2 | β | β | β | high |
35β | Z GLM-5.2 Zhipu AIπ¨π³1moOSS | 78.3 | 72.9 | β | 89.3 | β | β | β | $3 | β | β | medium |
36β | X Grok-3 xAIπΊπΈ1y | 78.2 | β | 75.3 | 80.9 | β | β | 77.0 | $15 | β | β | low |
37β | GPT-5 Medium OpenAIπΊπΈ1y | 78.0 | β | 68.3 | 85.3 | β | β | β | β | β | β | low |
38β | Q Qwen3.7 Max Alibaba Cloud / Qwen Teamπ¨π³2mo | 77.5 | 77.5 | β | 80.3 | 76.5 | β | β | $4 | β | β | medium |
39β | X MiMo-V2-Pro Xiaomiπ¨π³4mo | 77.4 | 77.4 | β | β | β | β | β | β | β | β | low |
40β | Z GLM-5V-Turbo Zhipu AIπ¨π³4mo | 77.3 | β | β | β | β | β | β | β | β | β | low |
41β | GPT-5.2 Pro OpenAIπΊπΈ8mo | 77.0 | β | 84.2 | 72.1 | β | β | β | β | β | β | medium |
42β | ERNIE 5.0 Baiduπ¨π³6mo | 76.8 | β | 65.0 | 73.9 | 92.5 | β | β | β | β | β | medium |
43β | X Grok-3 Mini xAIπΊπΈ1y | 76.4 | β | 71.5 | 80.1 | β | β | β | β | β | β | low |
44β | GPT-5.6 Luna OpenAIπΊπΈ1mo | 76.2 | 75.8 | 95.5 | 89.6 | 40.9 | β | 77.4 | $1 | β | β | medium |
45β | GPT-5.1 OpenAIπΊπΈ9mo | 76.0 | 74.2 | 57.1 | 85.3 | β | β | 89.6 | $10 | β | β | medium |
46β | GPT-5.1 Instant OpenAIπΊπΈ9mo | 76.0 | 74.2 | 57.1 | 85.3 | β | β | 89.6 | $10 | β | β | medium |
47β | GPT-5.1 Thinking OpenAIπΊπΈ9mo | 76.0 | 74.2 | 57.1 | 85.3 | β | β | 89.6 | β | β | β | medium |
48β | Claude Sonnet 4.6 AnthropicπΊπΈ5mo | 76.0 | 80.2 | β | 80.4 | 73.2 | β | 71.0 | $15 | β | β | high |
49β | o1 OpenAIπΊπΈ1y | 75.4 | 48.6 | 94.2 | 79.9 | 86.7 | β | β | $60 | β | β | high |
50β | Claude Opus 4 AnthropicπΊπΈ1y | 75.2 | 58.2 | 75.5 | 79.9 | 96.0 | β | 68.1 | $75 | β | β | high |
51β | B Seed 2.0 Lite ByteDanceπ¨π³6mo | 75.0 | 68.4 | 74.9 | 81.6 | β | β | β | β | β | β | low |
52β | M Kimi K2-Thinking-0905 Moonshot AIπ¨π³11moOSS | 74.8 | 67.3 | 84.2 | 83.4 | β | β | β | β | β | β | medium |
53β | T Hy3 Tencentπ¨π³1moOSS | 73.3 | 63.4 | β | 87.8 | β | β | β | β | β | β | medium |
54β | Midjourney v6.1 MidjourneyπΊπΈ2y | 73.0 | β | β | β | β | β | β | β | β | β | low |
55β | M Kimi K2.6 Moonshot AIπ¨π³3moOSS | 72.9 | 72.2 | β | 74.4 | β | β | 81.4 | $4 | β | β | high |
56β | S Step-3.5-Flash StepFunπ¨π³6moOSS | 72.9 | 70.3 | 80.9 | β | β | β | β | $0.40 | β | β | low |
57β | GPT-5.2 OpenAIπΊπΈ8mo | 72.8 | 80.9 | 71.4 | 69.7 | 74.2 | β | 81.9 | $14 | β | β | high |
58β | Z GLM-5 Zhipu AIπ¨π³6moOSS | 72.8 | 77.0 | β | β | β | β | β | $3 | β | β | low |
59β | DeepSeek-V4-Pro-Max DeepSeekπ¨π³3moOSS | 72.7 | 58.5 | β | 84.9 | 78.1 | β | β | $3 | β | β | high |
60β | o1-pro OpenAIπΊπΈ1y | 72.3 | β | β | 72.3 | β | β | β | β | β | β | low |
61β | Gemini 3.1 Pro GoogleπΊπΈ5mo | 72.3 | 68.8 | β | 89.3 | 52.3 | β | 81.7 | $15 | β | β | high |
62β | Gemini 3 Pro GoogleπΊπΈ8mo | 72.2 | 74 | 84.2 | 65.7 | 64.5 | β | 75.3 | β | β | β | high |
63β | GPT-5.5 OpenAIπΊπΈ3mo | 72.1 | 53.6 | 51.0 | 90.9 | 83.6 | β | 86.3 | $30 | β | β | high |
64β | GPT OSS 20B High OpenAIπΊπΈ1yOSS | 71.9 | β | 82.7 | 63.8 | β | β | β | β | β | β | low |
65β | Claude Opus 4.5 AnthropicπΊπΈ8mo | 71.8 | 82.3 | β | 61.3 | 78.0 | β | β | β | β | β | medium |
66β | Q Qwen3.7-Plus Alibaba Cloud / Qwen Teamπ¨π³2mo | 71.3 | 67.8 | β | 72.5 | 72.2 | β | 81.3 | $1 | β | β | high |
67β | X MiMo-V2-Omni Xiaomiπ¨π³4mo | 71.2 | 71.2 | β | β | β | β | β | β | β | β | low |
68β | M Kimi K2.5 Moonshot AIπ¨π³6moOSS | 71.1 | 52.5 | 79.3 | 84.9 | β | β | 69.2 | β | β | β | high |
69β | S Step3-VL-10B StepFunπ¨π³7moOSS | 70.6 | β | 66.3 | β | β | β | 77.2 | β | β | β | low |
70β | GPT-5.5 Instant OpenAIπΊπΈ3mo | 70.6 | β | 54.4 | 82.2 | β | β | 71.6 | $30 | β | β | medium |
71β | GPT-5 Codex OpenAIπΊπΈ11mo | 70.5 | 70.5 | β | β | β | β | β | β | β | β | low |
72β | Claude Opus 4.1 AnthropicπΊπΈ1y | 70.3 | 75.0 | 48.4 | 75.4 | 73.8 | β | β | β | β | β | medium |
73β | OpenAI Codex OpenAIπΊπΈ1y | 69.7 | 58.4 | β | β | β | β | β | β | β | β | low |
74β | Q Qwen3.5-122B-A10B Alibaba Cloud / Qwen Teamπ¨π³5moOSS | 69.6 | 65.1 | β | 82.4 | 63.7 | β | 72.2 | β | β | β | high |
75β | Gemini 3 Flash GoogleπΊπΈ8mo | 69.4 | 77.4 | 83.9 | 65.1 | 62.2 | β | 71.7 | $3 | β | β | high |
76β | MiniMax M3 MiniMaxπ¨π³2moOSS | 69.4 | 68.8 | β | β | β | β | 76.8 | $1 | β | β | medium |
77β | GPT-5.1 Codex OpenAIπΊπΈ8mo | 68.9 | 68.9 | β | β | β | β | β | β | β | β | low |
78β | Q Qwen3.5-27B Alibaba Cloud / Qwen Teamπ¨π³5moOSS | 68.8 | 66.0 | β | 82.4 | 60.6 | β | 71.9 | $2 | β | β | high |
79β | Claude Sonnet 4 AnthropicπΊπΈ1y | 68.7 | 54.1 | 66.8 | 71.2 | 90.0 | β | 63.7 | $15 | β | β | high |
80β | U Solar Pro 4 Upstageπ°π·1w | 68.4 | 61.9 | β | 86.3 | β | β | β | $1 | β | β | low |
81β | X Grok 4 Fast xAIπΊπΈ11mo | 68.2 | β | 73.4 | 54.4 | 98.7 | β | β | β | β | β | medium |
82β | GPT-5.4 OpenAIπΊπΈ5mo | 68.1 | 48.3 | 69.1 | 81.7 | β | β | 83.0 | $15 | β | β | high |
83β | GPT OSS 120B High OpenAIπΊπΈ1yOSS | 68.0 | β | 74.1 | 75.4 | 52.1 | β | β | β | β | β | low |
84β | M Kimi K2.7 Code Moonshot AIπ¨π³2moOSS | 67.6 | β | β | β | β | β | β | $4 | β | β | low |
85β | MAI-Thinking-1 MicrosoftπΊπΈ2mo | 67.6 | 45.2 | 80.5 | 80.3 | β | β | β | β | β | β | medium |
86β | Gemma 4 31B GoogleπΊπΈ4moOSS | 67.4 | β | β | 59.9 | 72.9 | β | 74.1 | $0.38 | β | β | medium |
87β | DeepSeek-V3.2 DeepSeekπ¨π³8moOSS | 67.4 | 67.6 | 75.0 | 73.8 | β | β | β | β | β | β | medium |
88β | L K-EXAONE-236B-A23B LG AI Researchπ°π·7mo | 67.2 | β | 74.6 | β | 59.8 | β | β | β | β | β | low |
89β | Q Qwen3.5-397B-A17B Alibaba Cloud / Qwen Teamπ¨π³5moOSS | 67.1 | 74.4 | β | 64.9 | 70.4 | β | β | β | β | β | medium |
90β | Nova 2 Pro AmazonπΊπΈ8mo | 66.9 | 67.0 | 73.8 | 76.2 | β | β | 37.9 | β | β | β | medium |
91β | Z GLM-4.7 Zhipu AIπ¨π³7moOSS | 66.9 | 56.8 | 78.8 | 78.0 | β | β | β | β | β | β | medium |
92β | Z GLM-5.1 Zhipu AIπ¨π³4moOSS | 66.8 | 52.4 | β | 85.3 | β | β | β | $4 | β | β | medium |
93β | M Kimi K2 0905 Moonshot AIπ¨π³11mo | 66.7 | β | β | 66.7 | β | β | β | β | β | β | low |
94β | Q Qwen3-Next-80B-A3B-Thinking Alibaba Cloud / Qwen Teamπ¨π³11moOSS | 66.5 | β | 66.4 | 69.3 | β | β | β | β | β | β | low |
95β | MAI-Code-1.1-Flash MicrosoftπΊπΈ3d | 66.5 | 66.5 | β | β | β | β | β | $1 | β | β | low |
96β | M Muse Spark MetaπΊπΈ4mo | 66.5 | 48.3 | β | 75.5 | β | β | 84.8 | β | β | β | high |
97β | Llama 3.1 405B MetaπΊπΈ2yOSS | 66.3 | 69.3 | 72.0 | 48.6 | 67.9 | 87.7 | β | $3 | β | β | high |
98β | Gemini 2.0 Flash Thinking GoogleπΊπΈ1y | 66.2 | β | β | 63.8 | β | β | 71.1 | β | β | β | low |
99β | Q Qwen3.6 Plus Alibaba Cloud / Qwen Teamπ¨π³4mo | 65.5 | 60.3 | β | 66.0 | 73.8 | β | 75.3 | $3 | β | β | high |
100β | Q Qwen3.5-35B-A3B Alibaba Cloud / Qwen Teamπ¨π³5moOSS | 65.2 | 58.7 | β | 80.8 | 57.8 | β | 69.2 | β | β | β | high |
Yes. Click any model to see its full benchmark profile, or use the comparison feature to compare up to 4 models side by side with radar charts and benchmark-by-benchmark scoring.