Neura Intelligence Index
Code generation benchmark — 164 Python programming problems measuring functional correctness
| # | Model | Score | Provider | Source |
|---|---|---|---|---|
| 1 | 95.1 | Anthropic | ||
| 2 | 93.8 | Anthropic | ||
| 3 | 92.4 | OpenAI | ||
| 4 | 92 | Mistral | ||
| 5 | 92 | Anthropic | ||
| 6 | 90.2 | OpenAI | ||
| 7 | 89 | Meta | ||
| 8 | 88.4 | Meta | ||
| 9 | 88.1 | Anthropic | ||
| 10 | 87.6 | OpenAI | ||
| 11 | 87.2 | OpenAI | ||
| 12 | 86.6 | Alibaba | ||
| 13 | 84.9 | Anthropic | ||
| 14 | 80.5 | Meta | ||
| 15 | 79.3 | Mistral | ||
| 16 | 73.2 | Cohere | ||
| 17 | 73.2 | NVIDIA | ||
| 18 | 72.6 | Meta | ||
| 19 | 62.2 | Microsoft |