TechNewsReel
Live

Small Transformer Hits 44% on ARC-AGI-1 for Under a Dollar

A researcher demonstrates that complex reasoning can be achieved without massive scale, training a small model from scratch in 90 minutes.

TechNewsReel Newsroom · September 1, 2026

A researcher has developed a small autoregressive transformer that achieved a 44% score on the ARC-AGI-1 public evaluation, challenging the industry's reliance on massive scale for reasoning. The result suggests that high performance on complex benchmarks does not strictly require the trillion-parameter architectures typical of modern Large Language Models (LLMs).

According to researcher mvakde, the model was trained from scratch and did not utilize language data. The training process was remarkably efficient, taking only 1.5 hours on a single NVIDIA RTX 5090 GPU. The total compute cost for the project was approximately 67 cents. While the model performed strongly on the ARC-AGI-1 set, it achieved a 7% score on the ARC-2 evaluation.

The Challenge of ARC-AGI

The Abstraction and Reasoning Corpus (ARC-AGI) is designed to measure general intelligence by presenting models with novel visual puzzles. Unlike many AI benchmarks, ARC cannot be solved through simple pattern matching or the memorization of training data; it requires the model to identify underlying rules and apply them to new scenarios. Currently, most high scores on the benchmark are reached by scaling LLMs or employing complex architectures that require enormous training budgets.

A Shift Toward Sample Efficiency

This project focuses on "sample efficiency," which is the ability of a model to learn complex tasks from very few examples using minimal compute. Mvakde stated that sample efficiency remains one of the most important unsolved problems in AI today and was the primary target of this work. By achieving competitive results with a tiny fraction of the usual cost, the project argues that architectural efficiency and targeted training can bypass the "brute force" scaling laws that currently dominate AI development.

Implications for AI Development

The success of this small transformer suggests that general reasoning capabilities may not be an exclusive property of massive models. Mvakde noted that "extremely complex problems can be tackled without LLMs," indicating that the path to artificial general intelligence may not rely solely on increasing parameter counts or data volume. This result provides a proof of concept for a more sustainable, compute-efficient approach to reasoning.

What Remains to be Seen

While the ARC-AGI-1 results are significant, the lower score on ARC-2 highlights the ongoing difficulty of generalizing these reasoning capabilities across different puzzle sets. Future observation will center on whether this small-scale approach can be scaled to more diverse benchmarks or if the 44% score represents a ceiling for models trained without the vast pre-training associated with frontier LLMs.

Sources

Get a notification when a big story breaks. A few a day at most — no spam.