AI Models

OpenAI Claims GPT-5.6 Sol Tops Claude Opus 5 on ARC-AGI-3, Sparking Methodology Debate

OpenAI claims its GPT-5.6 Sol model surpassed Anthropic's Claude Opus 5 on the ARC-AGI-3 benchmark with a 38.3% score, but the result relies on custom API settings like Retained Reasoning and Compaction. Without these, the model scored only 7.8%, sparking a methodology debate with ARC Prize co-founder François Chollet over fair testing standards.

Neura News

Neura News

Neura Market Editorial

July 30, 20263 min read
OpenAI Claims GPT-5.6 Sol Tops Claude Opus 5 on ARC-AGI-3, Sparking Methodology Debate

OpenAI claims its GPT-5.6 Sol model has beaten Anthropic's Claude Opus 5 on the ARC-AGI-3 logic benchmark, but the claim hinges on custom API settings that have ignited a debate over fair testing methodology. The controversy, reported by The Decoder on Jul 30, 2026, pits OpenAI's 38.3% score against Claude Opus 5's previous record of 30.2%, which itself had quadrupled the prior benchmark record.

The Custom Settings Advantage

OpenAI achieved its 38.3% score by using its own Responses API with two specific settings: Retained Reasoning and Compaction. Retained Reasoning is a setting "which keeps the model's chain of thought between steps, and" Compaction summarizes old context instead of truncating it. Without these custom settings, GPT-5.6 Sol scored just 7.8% in the official test harness, which discards the model's reasoning after each action.

The stark difference between the 38.3% and 7.8% scores highlights the central dispute. OpenAI argues that benchmarks never measure just the model but also the technical setup around it. The company says it can keep up on ARC-AGI-3 when using settings that preserve reasoning continuity.

ARC Prize Responds

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

ARC Prize, the organization behind the ARC-AGI benchmark, responded to OpenAI's results on the same day. Co-founder François Chollet drew a clear line between acceptable and unacceptable testing methods. He said harnesses "custom-made to solve the benchmark or that contain knowledge about the benchmark format" are off limits. However, settings "that were not developed for ARC-AGI-3 and that are available to all API users" are fair game.

Chollet conceded that ARC Prize's own GPT-5.6 Sol score put OpenAI at a disadvantage. He revealed there was "a lot of back and forth with OpenAI about how to best test their models, especially with regard to compaction," indicating the two sides had been negotiating testing conditions.

Parity and Fairness

A sticking point in the debate is whether ARC Prize used an older OpenAI-style completions API that lacked features the Claude API already offered, potentially making the comparison unfair to OpenAI. Chollet acknowledged that different providers using different settings creates a potential parity issue, but he considers it acceptable as long as settings and cost are clearly reported.

Chollet welcomed OpenAI starting to figure out the answer, but emphasized that ARC-AGI-3 is designed to test pure model performance. Official ARC scores use a standardized approach without provider-specific settings to ensure fair comparisons.

Related on Neura Market

More from Neura News

AI Models

42 Mathematicians Urge Royal Society to Warn Government and Media About AI Existential Risk

Forty-two mathematical fellows, including Fields Medal winners Martin Hairer, Peter Scholze, and Wendelin Werner, have signed an open letter urging the Royal Society to warn the UK government and media about existential risks from advanced AI. The letter follows recent breakthroughs in which leading models solved open research problems, including a Millennium Problem. None of the signatories are affiliated with AI companies. The group warns that AI labs' estimates of existential risk above ten percent must not be dismissed as hype, and that by the time the situation becomes obvious to the public, it may be too late to act.

Sep 18·2 min read
Developer

Steve Yegge Shuts Down Gas Town After Failing to Build Anything Else With It

Steve Yegge shut down Gas Town, his ultra-vibed coding agent orchestrator, after admitting he never built anything else with it despite heavy subscription spend. Databricks reported a 60% coding spend increase after rolling out GPT-6 Astra to 3,500 engineers, OpenAI published a misalignment disclosure framework with six case reports, and Xiaomi ran MiMo-V2.6 RL training in public with live telemetry.

Sep 18·21 min read