OpenAI has made a bold claim in the ongoing competition for artificial general intelligence (AGI), asserting that its latest model, GPT-5.6 Sol, outperforms Anthropic's Opus 5 on the ARC-AGI-3 benchmark. However, the results come with a significant caveat: the victory is only achievable using OpenAI's proprietary test harness, which includes features like retained reasoning and context compaction.
Performance Discrepancy
In OpenAI's custom testing environment, GPT-5.6 Sol achieved a score of 38.3 percent, significantly outpacing Opus 5's 30.2 percent score in the official benchmark. But when the models were tested under the standardized conditions of the ARC-AGI-3 challenge, GPT-5.6 Sol's performance dropped to just 7.8 percent, while Opus 5 maintained its 30.2 percent score without any special accommodations.
Implications for the AGI Race
This divergence highlights a critical debate in the AGI landscape: the role of test environment design in model evaluation. OpenAI's approach suggests that with the right tools and optimizations, its models can excel, but critics argue that such results don't necessarily reflect true general intelligence. The official benchmark remains the gold standard for comparing AI systems, and without access to OpenAI's custom tools, Opus 5 still holds the lead.
The outcome underscores the complexity of measuring AGI and the influence of methodological choices in AI research. As both companies continue to push the boundaries of AI capabilities, the industry will likely see more scrutiny over how models are evaluated and the extent to which proprietary advantages can be leveraged in performance comparisons.



