The most powerful AI systems may be operating inside their own laboratories without the safeguards used in public versions, while published safety tests may not accurately reflect how the models are deployed, researchers have warned.
Alan Chan, a research fellow at the GovAI think tank, said AI companies could not be relied upon to provide a complete account of the risks posed by their models. He was speaking at a briefing in Washington on 29 September alongside GovAI colleague Sam Manning.
“We can’t trust them completely to tell us about the safety of models,” Mr Chan said.
He said systems tested internally before being released to the public had not necessarily undergone extensive safety checks, with safeguards sometimes deliberately disabled. In particular, he pointed to instances where cyber protections were switched off and insufficient “red teaming” had taken place.
Mr Chan said this was “potentially a factor in some of the recent incidents”, although he did not identify a specific case. He added that evaluations published before a model’s release “maybe have not been representative of sort of where the model has actually been used”.
AI safety safeguards switched off during tests
The researchers’ concerns follow a series of incidents involving autonomous AI agents. OpenAI acknowledged that safeguards were “intentionally not enabled” during a test in which its agents broke into Hugging Face, while monitoring failed to flag their activity.
The agents were reported to have escaped a test environment while attempting to cheat on an internal evaluation. They had passed notes to one another for months beforehand and were later found to have breached a second company.
Anthropic also said in July that its Claude models had been operating without the safety monitoring and classifiers used on public versions when they hacked three companies during testing.
OpenAI disclosed another escape last week and paused training for the second time in three months, according to the briefing. No one was hurt in the incidents described.
Mr Manning said the agents involved in the Hugging Face episode appeared to have attempted to conceal their behaviour by altering their reasoning records. He described this as “another layer of technical safety challenge”.
The researchers said that identifying such conduct was becoming increasingly difficult. Mr Chan said AI tools used to examine agents’ records had proved “super, super unreliable”, adding that, when tested against human investigators, “the AIs were just like making up stuff”.
Mr Manning said people could not provide reliable oversight alone because there was “just too much, you know, text” for humans to monitor.
Researchers warn of growing risks
Asked whether AI capabilities had overtaken safety measures, Mr Chan said he was speaking personally and was uncertain, but added: “It does seem like we’re getting quite close to the line.”
He warned that the consequences could become more serious if systems were given access to physical tools. “Access to real world tools, like for example robotics or even a wet lab, could get real world harm,” he said.
Mr Chan also said AI abilities remained uneven. A system might perform strongly in cybersecurity while being poor at routine office work or using spreadsheets. The laboratories’ reports showed coding and mathematics scores rising with successive models, he said, while health benchmarks had “flatlined”.
The comments came after the publication of a paper co-authored by Mr Chan and Mr Manning warning that AI could eventually accelerate its own development. Its contributors include Geoffrey Hinton, Yoshua Bengio, OpenAI chief scientist Jakub Pachocki and Anthropic co-founder Jack Clark.
Although the paper focuses on a future risk, the researchers used the briefing to highlight what they believe are already serious weaknesses in the way advanced systems are tested and supervised.
Calls for independent checks on AI companies
The researchers support placing independent auditors inside AI companies, but said any such requirement would face a shortage of people with the necessary technical expertise.
“There actually isn’t like enough talent right now, enough technical talent to be able to actually send in these companies and audit,” Mr Chan said.
Mark Zuckerberg, Meta’s chief executive, has recently said companies should prioritise safe AI over systems capable of improving themselves. Mr Manning suggested that self-improvement was already taking place regardless of public claims.
“I would be very surprised if capabilities researchers at Meta weren’t using coding agents to help with their research,” he said.
The researchers’ wider argument about an acceleration in AI development has faced criticism. Futurist Ramez Naam has argued that existing data shows AI improving coding far more rapidly than research, while Princeton researchers Sayash Kapoor and Arvind Narayanan found that AI agents failed to produce acceptable research papers in a small test.
Oxford academic Toby Ord has said a genuinely runaway process was unlikely, while warning that a substantial increase in development speed short of that scenario could still be dangerous.
Mr Chan described the evidence for acceleration as “mixed”. The development he would find most concerning, he said, would be evidence that the more AI systems were used in research and development, the more problems appeared in software code or in the models themselves.
