OpenAI's GPT-6 Astra holds the highest general score on the Epoch Capabilities Index, though Anthropic's Claude Fable 5.1 leads when the measure narrows to software engineering, according to a Data Insight published September 16, 2026 by research organization Epoch AI.
The analysis, co-authored by Alexander Barry and Jaeho Lee, compares four recent models: GPT-6 Astra, Claude Fable 5.1, GPT-5.6 Sol, and Kimi K3. It uses the ECI, a composite measure of model capabilities, alongside two domain-specific variants, Math-ECI and SWE-ECI.
GPT-6 Astra's general ECI stands at 166.31, with a 90% confidence interval of 163.00 to 171.88, based on 16 benchmarks. Claude Fable 5.1 follows at 164.47, with a 90% confidence interval of 161.36 to 168.28, based on 15 benchmarks. GPT-5.6 Sol, also from OpenAI, sits at 161.81, with a 90% confidence interval of 159.57 to 165.20, based on 22 benchmarks. Moonshot's Kimi K3 records 157.63, with a 90% confidence interval of 155.38 to 160.45, based on 19 benchmarks.
The same order holds on math. GPT-6 Astra's Math-ECI of 169.83, with a 90% confidence interval of 165.98 to 175.37, rests on 4 benchmarks and sets a new record. Claude Fable 5.1 scores 165.73, with a 90% confidence interval of 162.84 to 175.91, also on 4 benchmarks. GPT-5.6 Sol reaches 162.76, with a 90% confidence interval of 159.85 to 166.24, on 4 benchmarks. Kimi K3 comes in at 158.39, with a 90% confidence interval of 155.23 to 162.76, on 4 benchmarks.
The math benchmark scores available for GPT-6 Astra are OTIS Mock AIME, FrontierMath Tiers 1-3, FrontierMath Tier 4, and ProofBench.
Software Engineering Reverses the Order
On software engineering, the ranking flips. Claude Fable 5.1's SWE-ECI of 167.44, with a 90% confidence interval of 161.95 to 178.53, is based on 3 benchmarks. GPT-6 Astra's SWE-ECI of 163.58, with a 90% confidence interval of 160.42 to 169.82, is based on 4 benchmarks. Kimi K3 takes third at 161.54, with a 90% confidence interval of 158.59 to 165.87, based on 5 benchmarks. GPT-5.6 Sol lands fourth at 160.47, with a 90% confidence interval of 157.94 to 165.85, based on 6 benchmarks.
The software engineering subset for GPT-6 Astra is MirrorCode, WeirdML, DeepSWE, and FrontierCode. MirrorCode, WeirdML, DeepSWE, and FrontierCode also feed the SWE-ECI calculations more broadly.
Domain-specific ECIs are refit on only one domain's benchmarks, measuring a model's strength relative to the general ECI. They retain the benchmark difficulty parameters from the general ECI fit and re-estimate each model's capability from a subset of benchmarks from a specific domain, which makes the resulting scores comparable to general ECI scores.
Launch Scores and the September Refreshes
The picture changed between the models' September 3, 2026 launch snapshots and the latest refresh. At launch, GPT-6 Astra had an ECI of 169.23, with a 90% confidence interval of 164.85 to 174.02, based on 9 benchmarks, with 1 SWE benchmark, MirrorCode. Claude Fable 5.1 at launch had an ECI of 162.88, with a 90% confidence interval of 160.10 to 166.34, based on 12 benchmarks, with 1 SWE benchmark, FrontierCode.
By the current snapshot dated September 15, 2026, GPT-6 Astra's ECI had settled at 166.31, with a 90% confidence interval of 163.00 to 171.88, based on 16 benchmarks, with 4 SWE benchmarks: MirrorCode, WeirdML, DeepSWE, and FrontierCode. Claude Fable 5.1's ECI had risen to 164.47, with a 90% confidence interval of 161.36 to 168.28, based on 15 benchmarks, with 3 SWE benchmarks: MirrorCode, WeirdML, and FrontierCode.
Stay ahead of the AI curve
The most important updates, news, and content — delivered weekly.
No spam. Unsubscribe anytime.
Between the September 3 and September 13 refreshes, GPT-6 Astra's ECI fell by 2.9 points and Claude Fable 5.1's ECI rose by 1.6 points. Based on pre-release evals, Epoch AI gave Astra an ECI of 169 when it launched, but this fell as more software engineering benchmark results became available. The ECI is refit whenever results are added, so scores move as data arrives.
The General, Math and SWE ECI scores are based on all benchmark scores available as of September 13, 2026.
Wide Intervals and a Relative Measure
The number of benchmark results behind each model's scores is small. Each domain-specific ECI rests on three to six results, so confidence intervals are fairly wide. GPT-6 Astra and Claude Fable 5.1 have overlapping 90% confidence intervals on all three ECI variants.
Domain-specific ECIs are relative measures. They show how a model performs in one domain relative to its overall ECI and do not track absolute progress within a domain.
Epoch AI provides the ECI as a composite measure of model capabilities and publishes other benchmarking tools. Its website includes sections for Data Insights, Publications, Data Explorers, and Benchmarks, among them the Epoch Capabilities Index, MirrorCode, FrontierMath: Open problems, and EBR-bench. Topics covered include AI progress, scaling, software progress, open models, capabilities and benchmarks, math, industry, leading companies, finances, geopolitics, infrastructure, chips, data centers, energy, impacts, adoption and use, economic impact, and the future of AI.
A Domain-specific ECI Explorer lets readers examine these scores and custom subsets.
Two data downloads accompany the analysis: General, Math and SWE ECI scores with 90% confidence intervals, in CSV, updated September 15, 2026; and GPT-6 Astra and Claude Fable 5.1 ECI at launch versus now, also CSV and updated September 15, 2026.
Epoch AI's work is free to use, distribute, and reproduce under a Creative Commons BY license with source and authors credited.
The Data Insight was published on September 16, 2026, three days after the latest refresh and one day after the current snapshot date.

