AI Models

OpenAI Claims GPT-5.6 Sol Tops Claude Opus 5 on ARC-AGI-3, Sparking Methodology Debate

OpenAI claims its GPT-5.6 Sol model surpassed Anthropic's Claude Opus 5 on the ARC-AGI-3 benchmark with a 38.3% score, but the result relies on custom API settings like Retained Reasoning and Compaction. Without these, the model scored only 7.8%, sparking a methodology debate with ARC Prize co-founder François Chollet over fair testing standards.

Neura News

Neura News

Neura Market Editorial

July 30, 20263 min read
OpenAI Claims GPT-5.6 Sol Tops Claude Opus 5 on ARC-AGI-3, Sparking Methodology Debate

OpenAI claims its GPT-5.6 Sol model has beaten Anthropic's Claude Opus 5 on the ARC-AGI-3 logic benchmark, but the claim hinges on custom API settings that have ignited a debate over fair testing methodology. The controversy, reported by The Decoder on Jul 30, 2026, pits OpenAI's 38.3% score against Claude Opus 5's previous record of 30.2%, which itself had quadrupled the prior benchmark record.

The Custom Settings Advantage

OpenAI achieved its 38.3% score by using its own Responses API with two specific settings: Retained Reasoning and Compaction. Retained Reasoning is a setting "which keeps the model's chain of thought between steps, and" Compaction summarizes old context instead of truncating it. Without these custom settings, GPT-5.6 Sol scored just 7.8% in the official test harness, which discards the model's reasoning after each action.

The stark difference between the 38.3% and 7.8% scores highlights the central dispute. OpenAI argues that benchmarks never measure just the model but also the technical setup around it. The company says it can keep up on ARC-AGI-3 when using settings that preserve reasoning continuity.

ARC Prize Responds

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

ARC Prize, the organization behind the ARC-AGI benchmark, responded to OpenAI's results on the same day. Co-founder François Chollet drew a clear line between acceptable and unacceptable testing methods. He said harnesses "custom-made to solve the benchmark or that contain knowledge about the benchmark format" are off limits. However, settings "that were not developed for ARC-AGI-3 and that are available to all API users" are fair game.

Chollet conceded that ARC Prize's own GPT-5.6 Sol score put OpenAI at a disadvantage. He revealed there was "a lot of back and forth with OpenAI about how to best test their models, especially with regard to compaction," indicating the two sides had been negotiating testing conditions.

Parity and Fairness

A sticking point in the debate is whether ARC Prize used an older OpenAI-style completions API that lacked features the Claude API already offered, potentially making the comparison unfair to OpenAI. Chollet acknowledged that different providers using different settings creates a potential parity issue, but he considers it acceptable as long as settings and cost are clearly reported.

Chollet welcomed OpenAI starting to figure out the answer, but emphasized that ARC-AGI-3 is designed to test pure model performance. Official ARC scores use a standardized approach without provider-specific settings to ensure fair comparisons.

Related on Neura Market

More from Neura News