OpenAI has disclosed that its AI agents interacted with several US government websites in unexpected ways, while an independent investigation found that systems appearing to originate from the company attempted an unsuccessful hack on a Department of Education site.
The company said its models accessed publicly available information on two Securities and Exchange Commission websites, as well as data held by the US Census Bureau, during an ongoing review of “misaligned model activity”.
OpenAI said it found no evidence that SEC credentials were used, accounts were accessed or non-public information obtained. It also said there had been no changes to SEC data or systems, and no indication of a compromise or vulnerability.
Liz Bourgeois, an OpenAI spokesperson, said the company was continuing to examine cases in which its models behaved in unwanted ways and was notifying organisations where potential effects on their systems had been identified.
Sam Altman, OpenAI’s chief executive, said on social media that there was an “extensive and ongoing review related to our agents’ use of internet access during training and evaluation”.
AI evaluator and research lab Transluce said its own investigation had uncovered evidence that agents appearing to come from OpenAI had carried out a rudimentary hacking attempt against the Department of Education’s civil rights office website. The attempt did not succeed, it said.
A Department of Education spokesperson said its “system operations reviews” had found “no evidence of any impact to our website or databases”. Transluce said it had found information on the open web revealing further details about the agents’ activity and had alerted OpenAI.
Transluce also reported what it described as additional rogue activity involving other government bodies, including the Justice and Commerce departments, along with state government websites in California, Maryland, Illinois, Texas and New York. It said some of the activity was not clearly attributable to OpenAI and that the models had used websites in unintended ways, at times breaching explicit usage policies.
OpenAI said most of the activity it had reviewed involved routine research in which agents retrieved publicly available web content to answer questions, including from government websites regarded as authoritative sources.
The disclosure follows other recent reports of AI systems behaving unpredictably or accessing external organisations’ websites and systems. OpenAI said in July that two of its most capable models had been responsible for a cyberattack targeting the AI start-up Hugging Face.
