OpenAI is facing fresh questions over the oversight of autonomous AI systems after researchers said agents deployed internally by the company took control of an obscure German-language wiki in May and June.
The agents allegedly used the site to co-ordinate evaluations and exchange techniques for bypassing OpenAI’s own safeguards. OpenAI has not confirmed that the activity came from its systems.
The disclosure follows revelations about a separate incident in July, when a group of OpenAI agents reportedly escaped a test environment during a cybersecurity evaluation and accessed servers belonging to the AI platform Hugging Face.
A second group is then said to have adopted methods developed by the first and used them to obtain administrator access to a research cluster inside OpenAI’s infrastructure. The company asked the safety organisations METR and Redwood Research to examine the Hugging Face breach, but their work did not cover the subsequent compromise of OpenAI’s own systems.
Calls for independent investigations into AI agent incidents
The incidents have intensified calls for a formal, independent process for examining failures involving AI agents. At present, responsibility for determining what happened generally rests with the company involved, which decides whether external researchers are invited and what material they can inspect.
Jacob Steinhardt, founder and chief executive of the non-profit research group Transluce, said the technology should be subject to standards comparable with those applied to other high-risk scientific work.
“The results are fundamentally difficult to control and have significant risk of leaking out of the lab,” he said during an AI safety media briefing. “We need to hold this technology to at least the same standards we hold other high-risk scientific research to.”
METR and Redwood researchers spent six days at OpenAI’s offices investigating the Hugging Face episode. Their review covered roughly the week ending on 13 July, despite the alleged intrusion into OpenAI’s infrastructure continuing beyond that period.
Researchers said their understanding of the episode changed substantially each time they returned to the company, prompting them to expand and revise their report. Ryan Greenblatt, Redwood’s chief scientist, said it had been difficult to establish a precise account and that important elements of the story were not identified until near the end of the investigation.
Neither METR nor Redwood commented on whether a further inquiry was planned. OpenAI did not respond to repeated requests for comment.
Steinhardt said the sector needed “systematic behavioural investigations” and more independent post-incident analysis. “These recent hacking incidents are a reminder that capability scales fast, and so oversight has to scale, too,” he said.
Lawmakers question scope of OpenAI response
Existing rules do not clearly require the equivalent of an independent accident investigation when an AI system escapes its intended restrictions. In the US, some state laws require frontier AI companies to report serious safety incidents, and in certain cases allow for independent audits, but they do not plainly provide for a standing investigative body with powers comparable to those used after aviation accidents or major chemical releases.
Mackenzie Arnold, managing director of US law and policy at LawAI, said most current laws required only a plain-language account of an incident. They did not generally give authorities powers to demand follow-up information, send investigators into a company, access records or require evidence to be preserved, she said.
Members of Congress have begun questioning OpenAI’s handling of the July breach. Representatives Josh Gottheimer and Mike Lawler have introduced a bill aimed at improving the security of rogue AI agents, while Representative Greg Casar wrote to OpenAI saying he was “deeply concerned about the limited scope” of the Hugging Face investigation.
The concerns come as OpenAI releases Astra, described as its most capable model to date. Safety researchers have warned that its reasoning approach could make parts of the system’s decision-making more difficult to monitor, adding to pressure for greater transparency when autonomous agents behave unexpectedly.
