ARC Prize published its own evaluation of GPT-6 Astra within hours of the model's launch on 3 September, and the headline number does not survive it. Run on ARC Prize's provider-neutral Standard harness, Astra scores 62.7% on ARC-AGI-3 Semi-Private. The 99.9% that led OpenAI's launch materials was produced by a different runner.

Two harnesses, two numbers

ARC Prize publishes both. The Standard harness lets a model carry forward notes it chooses to keep. The Provider Adapter harness, in ARC Prize's own words, "preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work." On the published results table the top Standard score is 62.7% at $26,098, at max reasoning effort. The 99.9% sits on the adapter side at high effort, at $18,817. At max effort the adapter returns 98.6% for $17,332.

The comparison that does not hold

This is the part that matters for every chart published on launch day. OpenAI's table put Astra's adapter score against Claude Opus 5 at 30.2% and GPT-5.6 Sol at 7.8% — figures measured on the Standard harness. Run all three the same way, the roughly 69-point gap over Anthropic's model becomes about 32 points. Still a lead, and a large one; not the one that was shown.

Cheaper as well as higher, which is the tell

The adapter run cost less than the standard one at every reasoning level — $18,817 against $40,705 at high effort. That rules out the usual explanation that a better score simply bought more compute. What the adapter adds is state reuse across calls, so part of what is being measured is OpenAI's inference plumbing rather than the model's reasoning. ARC Prize says so directly, noting its testers had no code interpreter or scratchpad, so results "should be understood as the combined performance of the model and its tools."

What the received framing gets wrong

"ARC-AGI-3 is saturated" circulated widely on 3 September. It is not. A third of the benchmark is unsolved under the provider-neutral runner. And the adapter is not a trick — preserving reasoning state is a real capability customers get. But it produces a system result, not a model result, and reporting the two as interchangeable is the error. A further detail undercuts the tidy scaling story: on the Standard harness the scores are non-monotonic. Reasoning effort set to "none" scores 35.2%, roughly double the 17.5% returned at "low".