Artificial intelligenceJuly 30, 2026· via The Decoder

OpenAI’s GPT-5.6 Sol outperforms rivals on ARC-AGI-3 under specific settings

OpenAI’s GPT-5.6 Sol outperforms rivals on ARC-AGI-3 under specific settings

Image : The Decoder

A fresh round of AI benchmarking has reignited debate over how fairly models are tested. OpenAI’s newest GPT-5.6 Sol variant has clocked a 38.3% success rate on the ARC-AGI-3 evaluation—more than four times higher than Anthropic’s Opus 5—but only when using OpenAI’s latest API and two undisclosed settings. Under the provider-neutral test setup, GPT-5.6 Sol managed just 7.8%, highlighting how implementation details can dramatically sway results.

The benchmarking dispute

ARC Prize, the organization behind ARC-AGI-3, markets its test as provider-neutral, designed to measure abstract reasoning across AI systems. Yet OpenAI’s internal testing suggests the official environment may rely on an outdated API, potentially handicapping newer models. The company argues that enabling its latest features and two additional configurations—neither described in public documentation—delivers a more accurate reflection of current capabilities. Anthropic has not yet publicly responded to the comparison.

What the numbers actually mean

The 38.3% score achieved under OpenAI’s bespoke conditions is still far from human-level performance, which typically hovers near 90% on ARC-AGI-3. Even so, the gap between OpenAI’s proprietary setup and the standardized test underscores a growing tension: benchmarks designed to be fair can become outdated as models evolve faster than their evaluation frameworks. Without transparent, consistently updated testing protocols, rankings risk reflecting tooling advantages rather than raw capability.

Why it matters

This episode reveals how fragile AI benchmarking remains when implementation details aren’t locked down. For developers, the lesson is clear: cutting-edge models can outperform rivals in real-world deployments while lagging in standardized tests—or vice versa. For users and regulators, it underscores the need for transparent, reproducible evaluation methods that keep pace with rapid model updates. The ARC-AGI-3 controversy isn’t just about scores; it’s about trust in the metrics that shape AI progress.


Source: The Decoder. AI-assisted editorial synthesis — TechnoExpress.

Read the original source on The Decoder →

← Back to home