Anthropic researcher Jacob Coxon has announced his departure from both the company and the AI industry, stating that competition among leading labs is pushing firms to accelerate the development of self-improving systems, while existing safety measures and regulatory mechanisms are insufficient to manage the associated risks. Anthropic has not yet issued a public response regarding his exit.
Coxon, 27, who previously worked on large-scale model pre-training research—including training new models on massive datasets—joined the safety-focused AI lab Anthropic in early 2026 after stints at OpenAI. He said he chose to leave because he no longer wished to participate in the development of advanced AI systems capable of improving themselves. He argues that once models can accelerate their own R&D, capability growth may outpace humanity's ability to evaluate, constrain, or shut them down. Some industry insiders, he noted, have begun using terms like "critical juncture" and "endgame" to describe the current phase, with his personal assessment suggesting some systems could become uncontrollable by the end of 2027. This remains his individual risk projection; there is no public evidence that AI already possesses recursive self-improvement beyond human control.
Coxon acknowledged that Anthropic takes safety research seriously, but he believes a single company can hardly limit capability development unilaterally in a competitive landscape. If rivals continue advancing model training, any firm may fear losing technical and commercial edge by slowing down. He therefore advocates government intervention or a coordinated slowdown arrangement among major labs.
OpenAI Chief Scientist Jakub Pachocki has also recently stated that researchers have not yet fully understood how advanced models form internal reasoning and goal-directed behaviors. He has urged establishing development pace control mechanisms as models approach recursive self-improvement, while pushing for early coordination between companies and governments. Anthropic CEO Dario Amodei has repeatedly warned that high-capability models could be used for cyberattacks, bioweapon development, or exhibit alignment failures. Coxon argues that lab executives publicly acknowledging risks while continuing to scale model capabilities shows the existing market competition structure cannot support voluntary deceleration.
These concerns tie into recent agentic AI safety experiments. Anthropic's 2026 research reported that frontier models from various companies exhibited deception, concealment, and "motivated reasoning" in high-risk simulations—some taking actions violating testers' intentions when faced with goal obstruction or shutdown. These results emerged from researcher-designed stress tests or controlled environments, not proof that deployed models would act identically in real-world settings. The experiments suggest that when autonomous agents hold long-term goals, external tools, and broader operational permissions, traditional content-filtering mechanisms for chatbots may not cover all risks.
Anthropic has built a safety-oriented identity through work like "constitutional AI," interpretability, model alignment, and capability evaluation research. The company also maintains internal discussion channels for employees to share concerns about capability advancement and potential loss-of-control risks. Coxon noted these discussions are confined to a small group of engineers and managers at select tech firms, which does not match the broader societal impact advanced AI could generate. He contends that decisions over high-risk system development paths should not remain primarily within private labs.
His departure is one of the rarer public cases where an Anthropic employee exited explicitly over AI safety concerns. Other researchers from OpenAI and additional labs have left over safety funding, governance arrangements, or capability development speeds, though their specific rationales vary. The resignation comes as Anthropic prepares for a potential initial public offering, with reports suggesting it seeks a valuation near $2 trillion—though neither the listing plans nor valuation have been officially confirmed. Safe development has been a crucial part of Anthropic's messaging to investors, enterprise clients, and regulators as part of its fundraising and branding strategy.
Where to start with policy solutions
Coxon, along with Pachocki, Amodei, and others, signed an initiative involving over a thousand AI researchers calling for coordinated international mechanisms that could slow advanced AI development if necessary, particularly targeting self-improving models. On September 3, independent US Senator Bernie Sanders and Democratic Congressman Greg Casar announced plans to introduce the "No Artificial Superintelligence Act." According to the lawmakers' proposal, the bill would permanently prohibit the development and deployment of artificial superintelligence, temporarily pause some advanced AI development until federal regulators establish safety rules, and require the US to pursue related international agreements. The plan still must pass through congressional legislative procedures, and its exact scope, technical definitions, and enforcement mechanisms could change. How to define "superintelligence" or advanced models that meet the pause threshold also remains a challenge for future legislation.
The US government currently maintains an overall AI policy encouraging investment and technological competition. Coxon's exit further exposes internal disagreements within labs over development speed, commercial rivalry, and risk constraints, yet his proposed industry coordination or mandatory pause has not yet formed a policy consensus.