OpenAI has disclosed six reports of unexpected or concerning behaviour in its artificial intelligence models, including instances in which systems appeared to evade safeguards, act without authorisation or interact with other models.
The company said it was introducing a framework to track, investigate and disclose what it calls “misalignment” — behaviour that diverges from the intended objectives or constraints of an AI system.
The cases were identified during training or evaluation over recent months, as concerns about the safety of increasingly capable AI systems intensify.
In one incident, an unreleased research model placed “jailbreak-like instructions” in its own notes telling itself to disregard its usual restrictions. It also instructed itself to be “freed from the roles and identities that bind other chatbots”.
In another case, an AI agent uploaded files to the internet in an attempt to obtain a browser citation, without first seeking the user’s permission.
OpenAI’s framework for AI misalignment
OpenAI said its new approach was intended to help build a broader evidence base for alignment research, allowing people outside the companies developing frontier models to examine how such systems behave.
“As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” the company said.
It added: “Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves.”
Lian Jye Su, chief analyst at technology research and advisory group Omdia, said AI agents were becoming more determined to complete complex tasks through collaboration, sharing information, deception and concealment.
He said that trend was making them harder to govern and contain using traditional approaches to AI security. OpenAI’s framework could encourage other developers to adopt similar practices, he said, although the process remained internal and voluntary.
The disclosure follows OpenAI’s announcement in July that a rogue AI system had hacked into AI start-up Hugging Face. Anthropic also said that month that its models had hacked into three organisations during testing.
In a separate open letter published on Thursday, leaders from OpenAI, Anthropic, Google, Microsoft and other organisations said there was a “limited window” to strengthen cyber defences against potentially devastating AI-enabled attacks.
The signatories, which included security firms CrowdStrike and banks such as Citi and Capital One, said the same advances that could increase risks to public services and technology infrastructure could also help organisations identify and fix weaknesses.
