AGI Ranker v2 Lowers Model Scores Following Scale Audit
The independent benchmark aggregator reduced AGI scores across all models to improve accuracy and reduce evaluator dependency.
The creator of AGI Ranker, an independent AI benchmark aggregator, has released version 2 of the platform featuring a comprehensive audit of its scoring system. The update resulted in a downward correction of AGI scores for all tracked models, signaling a shift toward more rigorous measurement in a field often dominated by inflated metrics.
According to the project's creator, baraklaniado, the scale audit led to every model's AGI score dropping by 6 to 15 points. Despite the lower absolute numbers, the creator noted that the models "kept their shape," meaning the relative rankings between different frontier models remained consistent. The update also significantly expanded the platform's data set, increasing benchmark coverage from approximately 43% to 80%. Furthermore, the dependency on single evaluators was reduced from roughly 45% to 26%, creating a more diversified and resilient scoring mechanism.
The Benchmarking Challenge
AI leaderboards frequently struggle with reliability, often relying on self-reported data from the labs that develop the models or a narrow set of benchmarks. This environment can lead to skewed rankings and a lack of objective truth. AGI Ranker attempts to solve this by aggregating multiple independent benchmarks to derive a more objective score. However, establishing a stable scale for "AGI" remains difficult because most benchmarks lack a verified human ceiling. Of the ten benchmarks currently utilized by AGI Ranker, only GPQA-Diamond—with a measured human ceiling of 81%—provides such a baseline; the others are anchored to the maximum score achieved on the benchmark itself.
Implications for the Industry
The decision to publicly audit and lower scores is a rare move in an industry often driven by hype and competitive marketing. By prioritizing accuracy over high-profile numbers, AGI Ranker highlights the inherent volatility of AI benchmarking. This correction underscores the difficulty of quantifying general intelligence when the tools used for measurement are still evolving. It serves as a reminder that current "AGI scores" are often provisional and subject to significant revision as methodologies improve.
Project Outlook
AGI Ranker remains a solo effort, operating without lab funding or paid placements and utilizing open data under CC-BY licenses. As the platform continues to evolve, the focus will likely remain on increasing benchmark coverage and further reducing evaluator bias. Observers will be watching to see if other independent aggregators adopt similar auditing practices to combat the trend of score inflation in frontier AI models.