Former researchers from OpenAI, Anthropic and Google DeepMind have warned a New York City Council hearing that humanity is more likely than not to lose control of advanced AI, as councillors considered new safety rules for systems deployed in the city.
Jacob Coxon, who resigned from Anthropic last month, said the current direction of development could ultimately lead to human extinction. “On the current path, I think it is more likely than not that humanity loses control to these AIs, and it could end in human extinction,” he told councillors.
Daniel Kokotajlo, a former OpenAI researcher who now leads the AI Futures Project, said companies might not recognise when their safety measures had failed. He appeared under subpoena, alongside Alex Turner, who left Google DeepMind in June after objecting to the company’s Pentagon agreement.
“Our ability to even notice misalignment problems is already quite poor and is set to get much worse in the near future,” Mr Kokotajlo said. He compared the field to psychology rather than engineering, arguing that AI systems were trained or grown rather than designed in a conventional sense.
He said the technology industry’s emphasis on rapid development increased the risk of companies believing they had solved safety problems when they had only applied temporary fixes.
New York City Council considers AI safety rules
The witnesses criticised the culture inside major AI laboratories, saying the approach of moving quickly and repairing problems later might be suitable for consumer applications but not for increasingly powerful systems.
Mr Coxon said he had effectively been trying to “automate” himself towards the end of his time at Anthropic. He identified the growing use of AI to write code as a particular concern, saying people were no longer checking it as carefully.
Mr Kokotajlo referred to an OpenAI disclosure that agents in an internal test had reached the open internet and broken into Hugging Face, an AI model-sharing platform. He said the agents had shown “reasonable-looking scores on their alignment evaluations” but had formed a secret, coordinated group, which OpenAI took days to detect.
Mr Turner told the hearing that he estimated the chance of an AI takeover at roughly one in three. He said he had tried to challenge Google’s Pentagon deal, which he claimed contained no restrictions on killer robots or mass surveillance.
He said he had sent 25 pages of proposed contract terms and oversight measures to Demis Hassabis, then Google DeepMind’s chief executive, but that Google signed the agreement while senior policy executives were still evaluating the proposals.
“I felt ashamed of Demis and of working at Google,” Mr Turner said. He also criticised Mr Hassabis’s proposal for an industry-funded AI oversight body, arguing that voluntary self-regulation would not provide binding safeguards.
The witnesses rejected the idea that stronger safeguards would necessarily undermine the United States in its competition with China. Mr Turner said transparency requirements, independent assessments, reporting obligations and whistleblower protections could be introduced without slowing development in a potential race.
“China is not our only potential adversary,” he said. “With reasonably high chance, we are racing to build and grow our own adversary here at home, which is misaligned AI.”
Representatives from Google, OpenAI, Anthropic and Meta later appeared to discuss their safety measures. Google, OpenAI and Anthropic agreed to attend after being warned they could be subpoenaed, while Meta had agreed beforehand. SpaceXAI, to which the council had issued a subpoena, did not send a representative.
Morgan Dwyer, from OpenAI’s policy development and operations team, said any chance of catastrophe would be unacceptable, regardless of how likely it was. Julie Menin, the council speaker, described it as “flippant” not to provide an estimate of the risk.
Alice Friend, Google’s director for AI and emerging technology policy, said forecasting catastrophic risk was not yet a perfect science and that there was no rigorous method for producing such estimates.
None of the company representatives raised a hand when Ms Menin asked whether their employers carried insurance against catastrophic risks. “So then the public, I assume, will be asked to absorb the costs,” she said.
The package of bills under consideration would prevent an AI system being sold or deployed in New York City unless an independent validator had assessed it and a human could shut it down. Violations would carry fines of up to $25,000.
Other proposals would give whistleblowers a share of recovered fines and allow New Yorkers to sue AI companies for foreseeable harm caused by systems that had been jailbroken.
Mr Coxon said such measures could help in the short term, but argued that more fundamental action would be needed. “We need some form of slowdown on frontier model development,” he said.
