Artificial Analysis published version 4.2 of its Intelligence Index on 4 September, and the rankings from it are already circulating as news: Claude Fable 5.1 first, GPT-6 Astra second with "a 4pt gain over GPT-5.6 Sol", Google seventh among labs. Those are not scores on the ruler anyone was quoting two days earlier.
What changed in the composite
The changelog, verbatim: "+ AA-Briefcase, our agentic knowledge work evaluation with a private test set + Surge's GDP.pdf, long context document reasoning across 4,592 PDF pages − GPQA Diamond, an exceptional scientific reasoning evaluation that has now been saturated … plus greater weighting on held-out test sets to prevent gaming". And: "40% of our Index weighting is now private, held-out test sets - double the figure from v4.1." The grading changed too: AA-LCR v1.1 "added a grading system prompt and corrected errors and ambiguities in answer keys", and for GDPval-AA v2 and AA-Briefcase, "we have improved our sampling and re-anchored the Elo scale".
The cleanest proof that the ruler moved
It comes from Artificial Analysis's own pages, two days apart. On 2 September: "Gemini 3.8 Flash scores 59 on the Artificial Analysis Intelligence Index … on par with sub-maximum reasoning efforts of GPT-5.6 Sol (xhigh, 59) and Grok 4.6 (medium, 59)." On 4 September, the live v4.2 leaderboard: Claude Fable 5.1 "scores the highest … with a score of 57, followed by GPT-6 Astra (max) with a score of 55". Nothing about Gemini changed in between. A mid-tier model scored 59 on Tuesday; on Friday the ceiling of the entire index is 57.
What the received framing gets wrong
"GPT-6 Astra shows a 4pt gain over GPT-5.6 Sol" is a v4.2-versus-v4.2 comparison being read as progress against a v4.1 baseline that no longer exists. Every before-and-after claim published this week — including any that pairs a launch-day number with a v4.2 rank — is comparing two different composites. A maintainer is entitled to retire a saturated evaluation; the problem is that the numbers on both sides of the change are being quoted in the same sentence.
The coincidence worth naming
The evaluation removed was the "exceptional scientific reasoning" one, and in the new composite Google falls to seventh among labs. That is not evidence of anything by itself, and the doubling of private held-out weighting is a genuine anti-gaming measure. But it is the kind of coincidence a benchmark maintainer should expect to be asked about, and the index is quoted in launch posts and procurement decks as though it were a stable scale.
