Artificial Analysis v4.2 Index Shifts to Private Tests to Combat Model Gaming
The updated Intelligence Index prioritizes agentic workflows and held-out datasets to better measure frontier model reasoning.
Artificial Analysis has released version 4.2 of its Intelligence Index, an interim update designed to keep pace with the rapid release cycle of frontier AI models. The update marks a strategic shift toward more complex, realistic tasks and a heavier reliance on private data to ensure benchmark integrity.
To prevent AI labs from "gaming" the results, the Index now weights private, held-out test sets at 40% of the total score. These private evaluations include AA-Omniscience, CritPt solutions, and the newly introduced AA-Briefcase, which evaluates agentic knowledge work. According to the data, Claude Fable 5.1 currently holds the top position in the overall Intelligence Index. GPT-6 Astra follows in second place, demonstrating a 4-point gain over GPT-5.6 Sol.
The Move Toward Agentic Evaluation
The v4.2 update accelerates several features originally planned for version 5, reflecting a broader industry trend to abandon saturated public benchmarks. Artificial Analysis has officially removed GPQA Diamond from the Index, citing the need for evaluations that mirror professional environments rather than static test sets.
Two primary benchmarks now drive the index's reasoning metrics. The AA-Briefcase evaluation focuses on agentic knowledge work, where Claude Fable 5.1 and Opus 5 currently lead. Meanwhile, the GDP.pdf test measures long-context reasoning through massive document synthesis. In this category, GPT-6 Astra leads with a 33.2% all-pass rate, followed by GPT-5.6 Sol at 28.2% and Claude Fable 5.1 at 26.2%.
Implications for the Frontier
This shift in methodology matters because traditional benchmarks are increasingly failing to differentiate between the top-tier models. By introducing multi-week project simulations and complex document analysis, the Index provides a more accurate proxy for actual productivity in professional settings.
Beyond raw intelligence, the update highlights operational efficiency. Artificial Analysis reports that GPT-6 Astra is more token efficient than nearly every other model currently operating at the intelligence frontier, suggesting that the next phase of competition will be defined by the balance of reasoning power and resource consumption.
What to Watch
As the industry moves toward "game-proof" evaluations, the reliance on private datasets will likely become the standard for third-party auditing. The transition from v4.2 to the full v5 release will be the next key milestone to watch, as it will likely further refine how agentic capabilities are quantified. For now, the lead remains a tight contest between Anthropic's Fable series and OpenAI's Astra and Sol models.