TechNewsReel
Live

MirrorCode Benchmark: AI Can Rebuild Complex Software in Hours

A new long-horizon benchmark from Epoch AI and METR reveals frontier models can compress weeks of human engineering into a single day.

TechNewsReel Newsroom · August 4, 2026

Epoch AI and METR have released MirrorCode, a rigorous new benchmark designed to test whether AI agents can autonomously reimplement entire software projects from scratch. Unlike traditional tests that focus on isolated bugs or single features, MirrorCode requires models to match a target program's output exactly on end-to-end tests without any access to the original source code.

The benchmark comprises 25 target programs spanning diverse technical domains, including cryptography, compression, Unix utilities, and bioinformatics. To ensure results reflect genuine reasoning rather than memorization, the researchers employed a "cheat-resistant" design. This includes sandboxing models to prevent internet access and utilizing held-out tests to stop models from simply mimicking lookup tables.

The Shift to Long-Horizon Coding

Most existing software engineering benchmarks evaluate short-term tasks. MirrorCode shifts this focus toward "long-horizon" coding, where an AI must manage a project's lifecycle end-to-end. This approach allows researchers to better understand the limits of autonomous agents and the potential risks associated with their increasing capabilities in software development.

Early results indicate that frontier models are already capable of significant autonomous work. Claude Opus 4.7, for example, achieved an overall score of 56% on the benchmark. In one standout performance, the model reimplemented "gotree"—a bioinformatics toolkit consisting of approximately 16,000 lines of Go and more than 40 commands. The AI completed the task in 14 hours at a cost of $251. By comparison, Epoch AI estimates that a human engineer working without AI assistance would take between two and 17 weeks to complete the same project.

Industry Implications

The ability of AI to rebuild multi-thousand-line projects based solely on observed behavior suggests a massive leap in productivity. Epoch AI believes a human engineer would take months to solve the most complex tasks within the MirrorCode suite. This capability indicates a fundamental shift in software engineering, where the time required to develop complex tools could be compressed from months to hours.

However, this acceleration introduces new systemic risks. The speed at which AI can generate functional, complex software means that malicious or misaligned code could potentially be produced at a scale and velocity that far outpaces current human oversight mechanisms.

Scaling Autonomy

MirrorCode is designed to accommodate the high inference budgets required for such deep work. In one of the benchmark's largest tasks, an AI agent worked for 19 days without any human intervention, incurring a cost of $2,600. As models continue to evolve, the focus will likely shift toward how these agents handle even larger codebases and whether they can maintain accuracy as the project horizon extends further.

Sources

Get a notification when a big story breaks. A few a day at most — no spam.