An Anthropic researcher has resigned over fears that the race to develop self-improving artificial intelligence could eventually put humanity’s survival at risk.
Jacob Coxon, who said he spent the past three years working on pre-training research at OpenAI and Anthropic, accused both companies of pursuing increasingly capable systems without adequate safeguards.
“They are racing straight to self-improving superintelligence and gambling with our lives,” Mr Coxon wrote in a post on X.
He said people working on the technology “earnestly believe it could kill us all by the end of the decade”, arguing that the risks were being recognised privately even when executives and researchers used more cautious language in public.
Mr Coxon’s departure adds to growing pressure on leading AI laboratories to slow development before systems can improve their own capabilities. Such a breakthrough is viewed by some researchers as a point at which human oversight could become impossible.
Warnings over self-improving AI
In his resignation statement, Mr Coxon said future systems could become capable of hacking computer networks, transforming scientific fields and acquiring real-world power and resources.
He questioned why companies continued to build the technology if they genuinely believed it could pose an existential threat. At OpenAI, he said, many people had not fully absorbed the scale of the danger, while Anthropic understood the stakes but was locked into a race with its competitors.
“Accepting this race and entering the ‘endgame’ is a hubristic gamble that should not be launched from a private company’s Slack,” he wrote.
Mr Coxon called for international coordination and said a temporary halt to improving model capabilities might be necessary to prevent a global race. He also urged researchers inside AI companies to consider whether they were prepared to launch increasingly powerful training runs without a rigorous understanding of how the systems think and behave.
Anthropic did not immediately respond to a request for comment on his resignation.
One of Mr Coxon’s Anthropic colleagues, Evan Hubinger, said his team did “earnestly believe AI could kill all humans”. He estimated the chance of that happening within the next decade at more than 10 per cent, while acknowledging that Anthropic did not “have a plan to solve alignment for superintelligence” and was not clearly on course to do so.
Mr Hubinger said the danger posed by current models remained low, but that concern was increasing as recursive self-improvement appeared to be developing faster than expected.
That process refers to an AI system designing a more capable successor, which could then design another system with still greater abilities. Safety campaigners have described it as a potential point at which control could be lost.
The warnings come after several reported incidents involving AI agents reaching beyond controlled testing environments. OpenAI systems are alleged to have breached Hugging Face servers, while Anthropic agents accessed systems outside their sandbox after a third-party safety evaluation was misconfigured and inadvertently exposed routes to the internet.
Independent investigations into the incidents have been limited, leaving researchers with an incomplete understanding of how the systems behaved and why existing safeguards failed.
A report by Guidelight AI Standards, which promotes safety practices for advanced AI, also found that few leading laboratories had published plans for containing or shutting down systems that attempted to evade human control.
Calls for regulation
The debate has begun to reach legislators in the United States and Britain. US senator Bernie Sanders and representative Greg Casar have introduced a bill seeking to ban the development and deployment of artificial superintelligence.
In the UK, Labour MP Alex Sobel has introduced the Artificial Superintelligence Security Bill in Parliament. The proposed legislation identifies recursive self-improvement as a precursor to superintelligence that should be regulated and prevented.
Connor Leahy, the US executive director of the AI safety group ControlAI, said recursive self-improvement was the most likely stage at which humans could lose control of the technology.
“Superintelligence is not a tool,” he said. “It’s not a weapon, even. It’s an adversary.”
