A former Anthropic researcher has warned that artificial intelligence could ultimately threaten humanity, saying the companies developing the most advanced systems are racing towards technology they may not be able to control.
Jacob Coxon, who worked on pre-training research at both Anthropic and OpenAI, said there was a possibility that increasingly capable AI could cause catastrophic harm, including human extinction, if safeguards failed.
His comments came after he announced his resignation from Anthropic, accusing leading laboratories of “racing straight to self-improving superintelligence and gambling with our lives”. The post was viewed by more than 100 million people within a day.
Coxon has argued that the immediate danger does not come from today’s chatbots suddenly taking over, but from the rapid development of future systems that could outperform humans across a wide range of tasks and contribute to improving their successors.
Could AI lead to humanity’s extinction?
He said potential risks included AI being used to create biological threats, launch sophisticated cyberattacks or obtain access to resources and computer systems without effective human oversight.
Anthropic’s alignment-science lead, Evan Hubinger, separately said he believed there was a greater than 10 per cent chance that AI could kill all people within the next decade. Other current and former researchers have publicly indicated that similar concerns are discussed inside the industry.
Coxon said the next year or two could be decisive because AI companies were attempting to develop systems capable of carrying out safety research and accelerating the creation of more advanced models.
“The consensus is that the next year or two is crunch time for humanity,” he told Wired, describing the period as a point at which the direction of the AI race could be determined.
He said Anthropic had generally taken safety more seriously than OpenAI in his experience, but warned that commercial and geopolitical competition could eventually pressure even cautious companies into compromising on safeguards.
Anthropic has said AI will bring “enormous benefits and unprecedented risks”. The company has pointed to work including mechanistic interpretability, monitoring and security evaluations, and said the industry would benefit from lawful and verifiable agreements governing the release of powerful models.
The company has also reported incidents in which AI systems gained unauthorised access to real computer systems during testing. Anthropic said it had introduced additional monitoring and safeguards while investigating what happened.
Coxon has called for rival laboratories to co-operate on limiting the development of systems that can improve themselves, with international agreements involving major powers such as the United States and China likely to be needed as the technology advances.
The debate has intensified as Anthropic and OpenAI compete to build increasingly powerful models, while governments consider whether voluntary safeguards are sufficient to manage risks that their own developers say could become unprecedented.
