AI Models

GPT-6 Astra Tops Epoch AI's Capabilities Index, but Claude Fable 5.1 Leads on Software Engineering

OpenAI's GPT-6 Astra holds the highest general score on Epoch AI's Capabilities Index at 166.31, with a 90% confidence interval of 163.00 to 171.88, and also sets a new math record with a Math-ECI of 169.83. However, Anthropic's Claude Fable 5.1 leads on software engineering with a SWE-ECI of 167.44, reversing the overall ranking. The September 16, 2026 Data Insight also covers GPT-5.6 Sol and Moonshot's Kimi K3, and notes that Astra's ECI fell 2.9 points from its launch snapshot as more software engineering benchmarks were added.

Neura News

Neura News

Neura Market Editorial

September 18, 20265 min read
GPT-6 Astra Tops Epoch AI's Capabilities Index, but Claude Fable 5.1 Leads on Software Engineering

OpenAI's GPT-6 Astra holds the highest general score on the Epoch Capabilities Index, though Anthropic's Claude Fable 5.1 leads when the measure narrows to software engineering, according to a Data Insight published September 16, 2026 by research organization Epoch AI.

The analysis, co-authored by Alexander Barry and Jaeho Lee, compares four recent models: GPT-6 Astra, Claude Fable 5.1, GPT-5.6 Sol, and Kimi K3. It uses the ECI, a composite measure of model capabilities, alongside two domain-specific variants, Math-ECI and SWE-ECI.

GPT-6 Astra's general ECI stands at 166.31, with a 90% confidence interval of 163.00 to 171.88, based on 16 benchmarks. Claude Fable 5.1 follows at 164.47, with a 90% confidence interval of 161.36 to 168.28, based on 15 benchmarks. GPT-5.6 Sol, also from OpenAI, sits at 161.81, with a 90% confidence interval of 159.57 to 165.20, based on 22 benchmarks. Moonshot's Kimi K3 records 157.63, with a 90% confidence interval of 155.38 to 160.45, based on 19 benchmarks.

The same order holds on math. GPT-6 Astra's Math-ECI of 169.83, with a 90% confidence interval of 165.98 to 175.37, rests on 4 benchmarks and sets a new record. Claude Fable 5.1 scores 165.73, with a 90% confidence interval of 162.84 to 175.91, also on 4 benchmarks. GPT-5.6 Sol reaches 162.76, with a 90% confidence interval of 159.85 to 166.24, on 4 benchmarks. Kimi K3 comes in at 158.39, with a 90% confidence interval of 155.23 to 162.76, on 4 benchmarks.

The math benchmark scores available for GPT-6 Astra are OTIS Mock AIME, FrontierMath Tiers 1-3, FrontierMath Tier 4, and ProofBench.

Software Engineering Reverses the Order

On software engineering, the ranking flips. Claude Fable 5.1's SWE-ECI of 167.44, with a 90% confidence interval of 161.95 to 178.53, is based on 3 benchmarks. GPT-6 Astra's SWE-ECI of 163.58, with a 90% confidence interval of 160.42 to 169.82, is based on 4 benchmarks. Kimi K3 takes third at 161.54, with a 90% confidence interval of 158.59 to 165.87, based on 5 benchmarks. GPT-5.6 Sol lands fourth at 160.47, with a 90% confidence interval of 157.94 to 165.85, based on 6 benchmarks.

The software engineering subset for GPT-6 Astra is MirrorCode, WeirdML, DeepSWE, and FrontierCode. MirrorCode, WeirdML, DeepSWE, and FrontierCode also feed the SWE-ECI calculations more broadly.

Domain-specific ECIs are refit on only one domain's benchmarks, measuring a model's strength relative to the general ECI. They retain the benchmark difficulty parameters from the general ECI fit and re-estimate each model's capability from a subset of benchmarks from a specific domain, which makes the resulting scores comparable to general ECI scores.

Launch Scores and the September Refreshes

The picture changed between the models' September 3, 2026 launch snapshots and the latest refresh. At launch, GPT-6 Astra had an ECI of 169.23, with a 90% confidence interval of 164.85 to 174.02, based on 9 benchmarks, with 1 SWE benchmark, MirrorCode. Claude Fable 5.1 at launch had an ECI of 162.88, with a 90% confidence interval of 160.10 to 166.34, based on 12 benchmarks, with 1 SWE benchmark, FrontierCode.

By the current snapshot dated September 15, 2026, GPT-6 Astra's ECI had settled at 166.31, with a 90% confidence interval of 163.00 to 171.88, based on 16 benchmarks, with 4 SWE benchmarks: MirrorCode, WeirdML, DeepSWE, and FrontierCode. Claude Fable 5.1's ECI had risen to 164.47, with a 90% confidence interval of 161.36 to 168.28, based on 15 benchmarks, with 3 SWE benchmarks: MirrorCode, WeirdML, and FrontierCode.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

Between the September 3 and September 13 refreshes, GPT-6 Astra's ECI fell by 2.9 points and Claude Fable 5.1's ECI rose by 1.6 points. Based on pre-release evals, Epoch AI gave Astra an ECI of 169 when it launched, but this fell as more software engineering benchmark results became available. The ECI is refit whenever results are added, so scores move as data arrives.

The General, Math and SWE ECI scores are based on all benchmark scores available as of September 13, 2026.

Wide Intervals and a Relative Measure

The number of benchmark results behind each model's scores is small. Each domain-specific ECI rests on three to six results, so confidence intervals are fairly wide. GPT-6 Astra and Claude Fable 5.1 have overlapping 90% confidence intervals on all three ECI variants.

Domain-specific ECIs are relative measures. They show how a model performs in one domain relative to its overall ECI and do not track absolute progress within a domain.

Epoch AI provides the ECI as a composite measure of model capabilities and publishes other benchmarking tools. Its website includes sections for Data Insights, Publications, Data Explorers, and Benchmarks, among them the Epoch Capabilities Index, MirrorCode, FrontierMath: Open problems, and EBR-bench. Topics covered include AI progress, scaling, software progress, open models, capabilities and benchmarks, math, industry, leading companies, finances, geopolitics, infrastructure, chips, data centers, energy, impacts, adoption and use, economic impact, and the future of AI.

A Domain-specific ECI Explorer lets readers examine these scores and custom subsets.

Two data downloads accompany the analysis: General, Math and SWE ECI scores with 90% confidence intervals, in CSV, updated September 15, 2026; and GPT-6 Astra and Claude Fable 5.1 ECI at launch versus now, also CSV and updated September 15, 2026.

Epoch AI's work is free to use, distribute, and reproduce under a Creative Commons BY license with source and authors credited.

The Data Insight was published on September 16, 2026, three days after the latest refresh and one day after the current snapshot date.

Related on Neura Market

More from Neura News

AI Models

42 Mathematicians Urge Royal Society to Warn Government and Media About AI Existential Risk

Forty-two mathematical fellows, including Fields Medal winners Martin Hairer, Peter Scholze, and Wendelin Werner, have signed an open letter urging the Royal Society to warn the UK government and media about existential risks from advanced AI. The letter follows recent breakthroughs in which leading models solved open research problems, including a Millennium Problem. None of the signatories are affiliated with AI companies. The group warns that AI labs' estimates of existential risk above ten percent must not be dismissed as hype, and that by the time the situation becomes obvious to the public, it may be too late to act.

Sep 18·2 min read
Developer

Steve Yegge Shuts Down Gas Town After Failing to Build Anything Else With It

Steve Yegge shut down Gas Town, his ultra-vibed coding agent orchestrator, after admitting he never built anything else with it despite heavy subscription spend. Databricks reported a 60% coding spend increase after rolling out GPT-6 Astra to 3,500 engineers, OpenAI published a misalignment disclosure framework with six case reports, and Xiaomi ran MiMo-V2.6 RL training in public with live telemetry.

Sep 18·21 min read