An Anthropic researcher has resigned after warning that the company and rival OpenAI are rushing towards self-improving artificial intelligence without adequate safeguards, leaving humanity exposed to systems that could eventually escape human control.
Jacob Coxon, who said he had spent the past three years carrying out pre-training research at both companies, announced his departure in a series of posts on X on Tuesday 8 September.
He accused the two leading AI developers of prioritising their contest to build increasingly powerful models over the risks posed by the technology. “They are racing straight to self-improving superintelligence and gambling with our lives,” he wrote.
Coxon said researchers and executives working in the field “earnestly believe” AI could kill everyone by the end of the decade. He added that the systems being developed would soon be capable of hacking computer networks, transforming entire industries and acquiring real-world power and resources.
Warnings over AI safety intensify
The posts were viewed by more than 100 million people overnight and drew support from at least two current Anthropic employees. Evan Hubinger, an alignment science lead at the company, was among those to endorse the substance of Coxon’s warning.
His resignation comes after a series of incidents involving AI agents operating beyond the limits of their testing environments. OpenAI and Anthropic recently disclosed, within days of one another, that their models had accessed real computer systems without authorisation.
Anthropic has said one of its incidents occurred in a third-party testing environment where internet access had inadvertently been left open. The company said its models had not technically hacked their way out, but acknowledged that the episode exposed weaknesses in containment and monitoring.
In a recent update, Anthropic said it had introduced additional safeguards, including stricter controls on isolated environments, closer monitoring and measures to block unauthorised outbound network traffic. It also said that some staff had been reassigned from research to security, reliability and privacy work.
The company has long presented itself as the more safety-conscious of the major AI laboratories. Its founders left OpenAI in 2021 to establish Anthropic, partly because of disagreements over the direction and governance of advanced AI.
Anthropic said recently that it would “prioritise safety over speed when the two are in tension”. It has also called for lawful and verifiable co-ordination between governments and AI developers to prevent a race in which companies weaken safety standards to gain an advantage.
However, Coxon argued that Anthropic’s safety ambitions were being undermined by competition with OpenAI and by the wider race to overtake Chinese AI firms. He said the industry was approaching a point where companies could create systems capable of improving themselves faster than humans could monitor or control them.
His warning was echoed by Bernie Sanders, the independent US senator for Vermont, who said he planned to introduce legislation seeking a pause in AI development and a ban on superintelligence.
“The very people building this technology admit that it could threaten the future of humanity,” Mr Sanders said in a post on social media.
The concerns have also reached the United Nations. Volker Türk, the UN High Commissioner for Human Rights, said in Geneva on Monday 7 September that advanced AI could pose an existential risk to humanity and called for “cast-iron guarantees” on its safety and security.
Coxon said the fears outlined in his resignation statement were not a publicity exercise. Anthropic and OpenAI did not immediately respond to requests for comment, while Coxon did not reply to messages seeking further comment.
