Tag
4 articles
OpenAI claims GPT-5.6 Sol outperforms Anthropic's Opus 5 on ARC-AGI-3, but only when using its own API features and additional settings not part of the official test.
OpenAI's GPT-5.6 achieved a tripling of scores on the ARC-AGI-3 benchmark after implementing just two API settings that enhanced reasoning retention and compaction.
Anthropic's Claude Opus 5 has achieved a 30.2% score on the ARC-AGI-3 benchmark, outperforming previous models like GPT-5.6 Sol and demonstrating advanced logical reasoning.
The ARC-AGI-3 benchmark challenges AI systems to match untrained human performance in interactive environments, with no frontier model achieving more than 1% success. The test strips away AI's typical advantages, exposing a gap in reasoning and adaptability.