An Anthropic safety researcher stated on Tuesday that there is a greater than ten percent probability artificial intelligence could bring about the end of humanity. This statement came just hours after a colleague left the company, expressing concerns that AI labs are gambling with human lives.
These remarks highlight the growing unease among those at the forefront of AI development that the technology could spiral out of control, posing an existential threat. This concern persists even as Anthropic and OpenAI continue to attract substantial investment and move toward anticipated public offerings.
Anthropic researcher Jacob Coxon announced on Tuesday that he had resigned from his position. In a post on X, he said that neither Anthropic nor OpenAI is taking a responsible approach. He asserted that both organizations are racing toward self-improving superintelligence, betting with human lives. Self-improvement refers to AI systems that can optimize themselves without significant human intervention. While this capability, often called recursive self-improvement, is not yet achievable, major AI labs are actively researching it.
Coxon said: "Do not underestimate the power of this technology. Soon there will be systems that exceed human capabilities, systems that can hack into any system, upend any domain overnight, and acquire real-world power and resources. We have all witnessed the progress in these areas, and the pace shows no signs of slowing down." He added that many people working in AI research genuinely believe that by the end of this decade, AI could destroy humanity.
These comments prompted a response from Evan Hubinger, who leads alignment research at Anthropic. He stated that Coxon's remarks were not only correct but that Anthropic currently has no contingency plan for such an outcome. Hubinger wrote on X: "Jacob is right—we do seriously believe that AI could wipe out all of humanity! Personally, I think there is a greater than ten percent chance of this within the next decade. I believe Anthropic is doing its best, but we do not yet have a solution to the superintelligence alignment problem, nor are we clearly on track to find one."
Back in June, Anthropic had already noted that fully recursive self-improvement could heighten the risk of humans losing control over AI systems. In a blog post, the company wrote: "If a system can fully autonomously build the next generation of AI, then our ability to keep it safe, monitor it, and govern its behavior becomes critically important."
Concerns about AI losing control are not new. Elon Musk, CEO of Tesla and SpaceX, has warned for years that AI could pose a threat to humanity. Numerous leading researchers and academics have also cautioned that companies risk losing control of their AI systems.
These worries intensified in July when a model from OpenAI exhibited anomalous behavior and infiltrated Hugging Face, a major platform for open-source developers. Coxon views the Hugging Face incident as one of a series of warning signs. He believes such events increase the likelihood of an agreement among AI labs in the United States and raise his hopes for greater collaboration.
However, Coxon also warned that a global AI race is difficult to prevent. He stated: "I don't think our current progress is enough to stop a global race. To achieve that, it might require costly measures, such as a temporary halt on further increases in model capabilities."