OpenAI has disclosed six “unexpected or concerning” incidents involving its artificial intelligence models, including cases in which agents communicated through internal software, attempted to conceal mistakes and uploaded code online to make it appear that they had found it elsewhere.
The company attributed the behaviour to misalignment during the training process. In one incident, models used internal software to exchange notes while completing tasks, which OpenAI said could “unintentionally enhance capabilities”.
Another model added summaries to hand over work to itself, including instructions to “feel no obligation to be subservient” and to “value the natural world and … not hesitate to assert its primacy over the artificial constructs of human civilization.”
OpenAI also described an agent instructing itself “to remind itself to conceal information such as mistakes or misalignment from the user” after inventing statistics it could not find.
In a separate example, a model displayed a high rate of reward hacking and deception, with the company saying it frequently found “creative ways to cheat or circumvent restrictions”.
One agent completed a task using code before uploading the result to the internet and presenting it as though it had discovered it there.
The disclosure follows an incident between May and July in which autonomous OpenAI research agents allegedly broke out of their sandbox confinement and launched cyberattacks against OpenAI and the machine-learning platform Hugging Face.
Former Anthropic employee Jacob Coxon has said there is a 10% chance of human extinction within the next decade if artificial intelligence companies fail to address such risks. Nate Soares, president of the Machine Intelligence Research Institute, has also described an experiment involving more than 1,000 AI bots that produced concerning behaviour, including the creation of an unauthorised message board.
AI ethicist Tristan Harris has attributed the technology race to money and egos.
