
OpenAI Claims GPT-5.6 Sol Tops Claude Opus 5 on ARC-AGI-3, Sparking Methodology Debate
OpenAI claims its GPT-5.6 Sol model surpassed Anthropic's Claude Opus 5 on the ARC-AGI-3 benchmark with a 38.3% score, but the result relies on custom API settings like Retained Reasoning and Compaction. Without these, the model scored only 7.8%, sparking a methodology debate with ARC Prize co-founder François Chollet over fair testing standards.










