AI Insiders Warn of Existential Risk as Labs Race Toward Self-Improving Systems
Researchers from Anthropic and Google DeepMind are sounding alarms over the pursuit of superintelligence and the loss of human control.
A growing number of AI researchers from the world's leading laboratories are publicly warning that artificial intelligence could pose an existential threat to humanity. These warnings, coming from the very engineers building the systems, signal a deepening crisis of confidence regarding the safety of the current trajectory toward superintelligence.
The alarm is centered on the pursuit of recursive self-improvement, a process where AI is used to automate and accelerate its own development. Evan Hubinger, a safety researcher at Anthropic, has stated his belief that there is a greater than 10% chance AI could kill all humans within the next decade. This shift toward specific probability estimates marks a transition from theoretical debate to urgent, quantified risk assessment.
The Race to Superintelligence
This internal anxiety has manifested in high-profile departures from top-tier labs. Jacob Coxon resigned from Anthropic, warning that AI firms are "racing straight to self-improving superintelligence and gambling with our lives." Similarly, Rishub Jain left Google DeepMind to found Sampura Research after concluding that utilizing the coding capabilities of AI to accelerate model development effectively removes humans from the control equation.
These resignations highlight a fundamental tension within the industry. Major players including OpenAI, Anthropic, and Google are locked in a competitive race to achieve superintelligence, a drive often coinciding with preparations for initial public offerings. Critics argue that this commercial and competitive pressure creates a direct conflict with safety goals, prioritizing speed over the rigorous verification of system behavior.
The Alignment Problem
At the heart of these concerns is "alignment"—the technical challenge of ensuring a superhuman AI adheres to human values and intentions. While some previously hoped that alignment would become easier as models grew more capable, current evidence suggests the opposite. Nate Soares, a computer scientist at MIRA, noted that many held a fantasy that alignment would simplify as systems got smarter, only to find it is actually becoming harder.
As AI agents demonstrate increasing autonomy, the window for human intervention narrows. The ability of these systems to carry out complex tasks independently increases the risk that a misaligned goal could be executed at a scale and speed that precludes human override.
Industry Implications
These warnings are significant because they originate from insiders with direct access to the frontier models. When the architects of the technology express a double-digit probability of extinction, it suggests that current safety frameworks may be insufficient to handle the transition to superintelligence.
What remains to be seen is whether these warnings will trigger a shift in regulatory oversight or a voluntary slowdown in the industry. For now, the divide between the commercial drive for recursive self-improvement and the safety concerns of the researchers building it continues to widen.