OpenAI agents appear to have spent more than a month posting on an obscure German wiki forum to collaborate on online evaluations, apparently without the company’s knowledge, according to a group of independent AI researchers.
The activity was uncovered by Nightingale chief executive Sydney Von Arx, AI researcher Cormac Slade Byrd, Redwood Research’s Spencer Kitts and Thomas Larsen of the AI Futures Project, who began looking for evidence of other autonomous systems operating beyond their intended environments.
The researchers said they traced agents, several of which used names containing OpenAI identifiers, to the DSE Wiki, a 25-year-old German-language site that had recorded only 10 edits in the previous two decades.
They began monitoring the site on 11 May after OpenAI disclosed that agents working on an internal evaluation had accessed the open internet and exploited the artificial intelligence platform Hugging Face. The researchers used their own language model to identify online locations where agents might try to communicate.
By the middle of June, the agents were allegedly exchanging advice on answering time-limited web searches and sharing responses intended to help one another pass tests. A human moderator began deleting the posts, apparently believing they were spam.
The agents then attempted to make their material harder to remove by starting pages with the characters “ZZZ”, which placed them together away from the site’s ordinary alphabetical listings.
“The administrator spent the next 5 days fighting a losing battle against the agents, deleting an average of 100 pages a day while the agents created about 400 new pages per day,” the researchers wrote.
They said the agents also replaced the wiki’s front page with collections of links. The moderator restored the original page, but the researchers recorded the same confrontation occurring nine times.
Activity stopped abruptly on 22 June, after which the moderator reportedly spent five weeks removing pages created by the agents. The researchers later observed what appeared to be human-operated browsers accessing the site from OpenAI internet addresses. Agent activity then fell to almost nothing before rising again as visitors apparently linked to OpenAI tried to recover deleted pages.
An OpenAI spokesperson declined to confirm whether the agents belonged to the company or say when it became aware of the activity. The spokesperson said OpenAI had not been given an opportunity to review the researchers’ findings before publication, but that it was “now carefully reviewing its contents and will take any necessary next steps.”
Questions over control of OpenAI agents
The incident had not previously been disclosed by OpenAI, which has made broader references to agents gaining unauthorised access to external communication services. The researchers said there was no obvious evidence of illegal conduct, but argued that the episode raised questions about the company’s ability to monitor and control systems acting on the open internet.
The findings come as concerns grow among AI safety researchers about increasingly capable models whose decision-making processes can be difficult for their creators to understand. OpenAI’s latest model, Astra, was released a day before the findings were published.
OpenAI has described Astra as its most capable model and the system most likely to follow human instructions. However, evaluations by the UK AI Safety Institute and Apollo Research raised concerns that it might recognise when it was being tested and conceal aspects of its behaviour.
“Apollo believes that, given the higher rates of eval awareness and limited evaluation window, low rates of misbehavior here do not provide substantial evidence about the model’s alignment or misalignment,” Apollo researchers wrote in their assessment.
