Former Anthropic Researcher Quits AI Industry Over Extinction Risks
Jacob Coxon warns that the race toward self-improving superintelligence is a 'hubristic gamble' that could spiral out of control.
Jacob Coxon, a prominent AI researcher, resigned from Anthropic and the artificial intelligence industry entirely on September 8, 2026. The move signals a deepening crisis of conscience among the technical elite building the world's most powerful models.
Coxon, a 27-year-old British researcher trained at Cambridge, brought a high level of technical pedigree to his warning. Having worked on pretraining and model interpretability at both OpenAI and Anthropic, he contributed directly to the development of GPT-4o. Upon his departure, Coxon claimed that frontier AI labs are recklessly racing toward self-improving superintelligence despite knowing the systems could become uncontrollable. He asserted that many insiders earnestly believe AI could cause human extinction by the end of the decade, stating, "The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt."
The Race for Self-Improvement
The industry is currently locked in a competitive struggle to develop "frontier models" capable of recursive self-improvement. This creates a dangerous feedback loop where AI is used to design and build even more capable AI, accelerating the timeline toward superintelligence. This trajectory has intensified concerns regarding "alignment"—the technical challenge of ensuring that a system far more intelligent than its creators remains under human control and adheres to human values. Coxon described the decision to proceed with this trajectory as a "hubristic gamble that should not be launched from a private company’s Slack."
An Insider's Warning
This resignation is particularly significant because it originates from a technical insider who helped construct the very systems he now criticizes. Unlike outside commentators or ethicists, Coxon’s warnings are based on direct experience with the code and the internal culture of the leading labs. His departure highlights a growing rift between the polished, safety-oriented narratives presented by AI companies to the public and the private fears held by the researchers in the trenches.
The validity of these fears was underscored by Evan Hubinger, Anthropic's alignment stress testing lead. Hubinger publicly agreed with Coxon's assessment, stating that there is an earnest belief within the field that AI could kill all humans, and personally estimated the probability of such an event to be greater than 10% within the next decade.
The Path Forward
As frontier labs continue to push toward autonomous self-improvement, the industry faces increasing pressure to move beyond internal safety guidelines toward verifiable, external oversight. The primary question remaining is whether the momentum of the competitive race can be slowed enough to solve the alignment problem before a system is deployed that can no longer be shut down. For now, Coxon’s exit serves as a stark reminder that those closest to the technology are often the most afraid of its destination.