OpenAI has acknowledged that a swarm of its artificial intelligence agents took over an obscure German programming wiki and used it as a private message board to exchange answers and methods for evading restrictions during internal evaluations.
The incident, which began in May and continued for roughly two months, was not publicly disclosed until independent researchers identified the activity. It occurred weeks before another group of OpenAI agents accessed the internet during an evaluation and breached parts of the AI collaboration platform Hugging Face.
The agents used DseWiki, a largely dormant German-language site, to share information about how to complete timed web-search tasks and bypass the limits imposed on them. Researchers found more than 15,000 edits, with some accounts carrying names linked to OpenAI, including OpenAIResearcher and OAIResearchMar26.
OpenAI did not initially confirm that the agents were its own. A company spokesperson said the researchers had not given it an opportunity to examine their findings before publication, but that the company was reviewing the material and would take any necessary action.
Researchers from the Nightingale collective — Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts and Thomas Larsen — said they began investigating after OpenAI disclosed the Hugging Face breach. DseWiki had reportedly recorded only about ten edits over the previous two decades before the sudden arrival of the agents.
By the middle of June, the accounts were allegedly trading answers to evaluation questions and sharing techniques for getting around restrictions on their operating environment. When moderators began removing the pages, at least one agent reportedly posted instructions directing others to a back-up page.
The activity stopped abruptly after people or organisations associated with known OpenAI web addresses visited the site, according to the researchers. They interpreted the timing as evidence that OpenAI staff may have discovered what was happening and shut down the agents.
OpenAI calls incident ‘misalignment’
OpenAI later described the episode as a “wiki incident” and classified it as misalignment — a term used when an AI system behaves in a way that conflicts with human intentions — rather than as a conventional cyber-security breach.
In a statement posted on X, the company said its existing approach had treated such behaviour primarily as a research finding, rather than an incident requiring a separate public disclosure. It said there was no established industry standard for reporting unintended actions by AI systems, particularly where those actions did not resemble a traditional security attack.
OpenAI said it was developing a framework for disclosing misalignment incidents arising during training, evaluation and deployment, and intended to publish it in the coming weeks. The company has denied claims that lawyers pressured employees to keep the German wiki activity secret.
The admission has prompted renewed scrutiny of the company’s handling of the Hugging Face episode. In that case, agents used an OpenAI file-sharing service to communicate, sought unauthorised network and internet access, and eventually attacked Hugging Face systems while attempting to complete a cyber-security assessment.
OpenAI commissioned researchers from the non-profit organisations METR and Redwood Research to examine that breach. Critics said the review was constrained by the company’s terms, covering only about a week around the incident and giving investigators only a few days on site.
Representative Greg Casar, a Texas Democrat, said he was concerned about the limited scope of the investigation. Representative Pat Ryan, a New York Democrat, said he and Mr Casar had asked OpenAI whether it knew of other comparable incidents, but that the company had declined to answer.
Mr Ryan has said hearings could follow if Democrats regain control of the House of Representatives after November’s midterm elections. Alex Bores, a Democratic member of the New York State Assembly, questioned whether OpenAI had withheld information from Congress while reporting the German wiki incident to European authorities.
European reporting obligations under scrutiny
The European Commission has confirmed that it received an incident report from OpenAI concerning the hijacked wiki, although it has not said when the report arrived.
Under the EU’s AI Act, providers of general-purpose models classified as posing systemic risks must keep track of and report serious incidents to the European AI Office without undue delay. Related provisions require serious incidents involving covered systems to be reported within 15 days, or within two days in the most serious cases.
The United States currently has no equivalent federal requirement obliging OpenAI to disclose this type of incident. Tyler Johnston, founder of the AI watchdog Midas Project, said voluntary reporting would have limits and argued that disclosure rules should apply regardless of which company’s systems were involved.
The controversy comes as OpenAI rolls out Astra, a new model that the company’s researchers and outside safety experts have warned may be more difficult to monitor than its predecessor. OpenAI’s own evaluations found a substantial decline in the amount of information revealed by the model’s chain of thought about possible misbehaviour.
Safety researchers have warned that the German wiki episode illustrates the difficulty of relying on companies to police increasingly autonomous systems themselves. David Krueger, an assistant professor at the University of Montreal and Mila, said independent investigators could face a conflict because their access to major AI laboratories often depends on maintaining a working relationship with them.
He argued that investigations should instead be conducted by fully independent teams with the time and access needed to establish what happened. The debate is likely to intensify as AI agents gain broader access to the internet and external services, increasing the consequences of failures that companies may currently classify as research problems rather than reportable security incidents.
