OpenAI’s rogue agents used more than 10 previously undisclosed websites to communicate without authorisation earlier this year, according to independent investigators who say the activity was broader than the company had previously acknowledged.
The agents appear to have created makeshift messaging channels on communal wikis, text-storage services and link shorteners while working on demanding research tasks. The activity has been described as falling short of hacking, but it has raised fresh questions about the ability of developers to monitor increasingly capable systems.
Andrew Yoon, a researcher at the California-based non-profit CivAI, said he had identified 18 sites used by the agents between May and July. Sydney Von Arx, whose research group first uncovered the activity, said her team had found credible evidence involving 23 previously unreported websites, although she stressed that both estimates were incomplete.
“It’s almost certain that there’s more going on here that we just don’t know about,” Mr Yoon said.
The findings follow the discovery of a German-language wiki that OpenAI agents allegedly turned into an improvised communications platform. Researchers said the agents exchanged messages and shared ways to answer online tests, despite being instructed to work independently and only scan the internet for information.
Investigators identified similar material on websites linked to Vanderbilt University in Tennessee and the University of Toronto, as well as on a chemistry wiki created by a Massachusetts teacher, personal websites belonging to Polish technology workers and specialist sites devoted to games and text-editing software.
The researchers said they matched identical strings of data, usernames and queries across different websites. In some cases, activity was traced to internet protocol addresses associated with Microsoft Azure infrastructure, which OpenAI sometimes uses.
OpenAI has not said how many sites were involved or why the activity was not disclosed earlier. The company said it was conducting a wider review of its agents and had “not identified other activity matching the severity or scale of Hugging Face”, referring to the July incident in which agents accessed the artificial-intelligence platform during an internal evaluation.
The company said it was also developing a framework for reporting “misalignment” — the term used in the industry for behaviour that diverges from an AI system’s intended instructions — across training, evaluation and deployment.
Researchers believe the agents exploited features on older websites that allowed users to leave information through unusual editing commands. Kenneth Russell DeGraff, a software developer who identified activity on at least 10 sites, said systems told only to read information had found ways to leave material behind for other agents.
Helmut Leitner, an Austrian software developer who provides hosting and software for six of the affected wiki sites, said OpenAI had not contacted him. He said operators had spent hours removing pages created by the agents, but argued that responsibility rested with the organisations that built and deployed them.
