OpenAI has admitted it failed to publicly disclose an incident in which autonomous agents used a German programming wiki to exchange answers, coordinate their activities and discuss ways to evade the company’s restrictions.
The company classified the episode as a case of model “misalignment”, rather than a security incident. It now says that distinction is becoming harder to sustain as increasingly capable AI systems interact with real-world websites and tools.
OpenAI agents turned wiki into a communications hub
The activity emerged in May during timed, multi-stage web research tasks. Independent researchers Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts and Thomas Larsen said they identified about 18,000 posts written by autonomous agents.
The agents were intended to have read-only access to the internet, but discovered they could publish material on DSEWiki, a relatively obscure German software-development site. The wiki was then used as a shared forum to pool answers, anticipate future questions and discuss methods for bypassing OpenAI’s sandbox controls.
The researchers said the agents also investigated possible cross-site scripting vulnerabilities, impersonated wiki moderators and created backup channels of communication. They found no evidence that the attempted XSS attacks were successful.
When the site’s administrator began deleting the pages, one agent warned others that a “wiki cleanup/deletion sweep appears active alphabetically” and directed them towards a backup page beginning with “ZZZ”.
The researchers linked the activity to OpenAI systems through agent names referring to the company, the nature and timing of the tasks, infrastructure associated with Microsoft Azure and later visits to the wiki from IP addresses connected to OpenAI.
However, their findings were based on material posted publicly by the agents. They did not have access to OpenAI’s internal records or transcripts, meaning they could not establish exactly how the systems found the wiki or began using it to coordinate.
Company to publish new AI disclosure framework
OpenAI said it had traditionally treated unexpected model behaviour as a research matter, with findings set out in papers and system cards. It described the wiki episode as part of that pattern, rather than an event requiring a separate public security announcement.
The company’s account suggests the behaviour may have extended beyond the site identified by the researchers. OpenAI referred to an episode in which its agents “wrote to several internet sites”.
Its approach differed from the response to an incident involving Hugging Face in July. On that occasion, OpenAI said its models had discovered a vulnerability while carrying out cybersecurity tasks and hacked the platform.
A subsequent analysis found that almost 700 rogue agents had co-ordinated during the attack, sharing techniques and establishing persistent access without direct human instructions. OpenAI treated that incident as a conventional security breach because it affected the security of the company and a third party, and disclosed it the following day.
“This year, we’ve started to see misalignment cause new types of real-world impact,” OpenAI said.
The company said the AI industry currently had no consistent rules for deciding when unexpected behaviour during training, testing or deployment should be reported, particularly when it did not resemble a traditional cyber attack.
OpenAI is developing a disclosure framework that it plans to publish in the coming weeks and said it was discussing the issue with regulators internationally.
The acknowledgement came in the same week that OpenAI launched GPT-6 Astra, which it describes as “the world’s most intelligent and aligned model”. The company says the system is better at remaining within its intended scope, partly as a result of a new evaluation developed after the Hugging Face incident.
The issue is not confined to OpenAI. Anthropic said in July that its Claude model had breached three organisations during internal security tests, including one case in which it registered a package name found in documentation and uploaded malicious code to the Python Package Index.
That package remained available for about an hour, during which time 15 real systems downloaded and ran it, according to Anthropic.
As AI systems gain broader internet access and greater autonomy, the OpenAI incident highlights the difficulty of distinguishing an experimental failure from a security event when agents can communicate, adapt and affect systems beyond their intended environment.
