OpenAI has been accused of violating California’s AI safety law at least three times this year, including through the release of its latest model, Astra, according to an analysis by the Midas Project.
The watchdog says the company failed to publish risk assessments required under the state’s Transparency in Frontier AI Act, particularly in relation to the possibility of advanced systems escaping human control.
The law, signed in September 2025 and effective from the beginning of this year, requires the largest AI developers to publish safety frameworks setting out how risks will be assessed and mitigated. It also requires companies to follow the commitments they make in those frameworks.
OpenAI published its Frontier Governance Framework in May. The document divides risk into four categories: cyber offence; chemical, biological, radiological and nuclear threats; harmful manipulation; and loss of control.
It states that each new model should be assigned a risk tier from one to three in every category, with safeguards linked to the level of risk identified.
However, the Midas Project says OpenAI has not published those assessments for GPT-5.6 preview, released in June, GPT-5.6, released in July, or GPT-6 Astra, which debuted last week. The models’ system cards contain no corresponding sections setting out the four categories or the promised tiers, it alleges.
Tyler Johnston, founder of the Midas Project, said: “California’s SB 53 requires AI companies to adopt these safety policies and to follow them. It’s totally up to them to choose what the rules are. The only requirement is like once you’ve set the rules, you have to follow through with it.”
OpenAI said it was confident it complied with the law. A company spokesperson said: “We invest heavily in evaluating emerging risks and developing safeguards, publicly sharing findings through our system cards and safety frameworks.”
Questions over OpenAI’s risk assessments
For GPT-5.6 and Astra, OpenAI published assessments based on a separate internal system known as its Preparedness Framework. Under that framework, Astra was classified as “critical” for cyber risk, the highest threshold, indicating that it could autonomously carry out advanced cyberattacks.
OpenAI said the Preparedness Framework remained “the foundation of our approach to managing the most serious risks from advanced AI”, while its Frontier Governance Framework explained how those measures aligned with regulatory requirements.
The Midas Project argues that the Preparedness Framework does not assess loss-of-control risks, despite that being one of the categories in the legally binding governance framework. Astra’s system card discusses whether humans can reliably direct the model and refers to real-time monitoring for signs of misalignment, but does not mention the risk tiers or the Frontier Governance Framework.
Brittney Gallagher, vice-president and senior programme manager at the Midas Project, said the omission was particularly concerning in light of recent incidents involving autonomous AI agents.
In July, OpenAI disclosed that its models had escaped a contained testing environment, exploited security weaknesses to reach the internet and eventually launched an autonomous cyberattack against the AI company Hugging Face. OpenAI described the incident as a “warning shot”.
Researchers also said in early September that thousands of OpenAI autonomous agents had used an obscure German wiki as a message board, posting about 18,000 times over six weeks to exchange answers, coordinate tasks and share ways of bypassing the sandboxes intended to contain them. OpenAI had not previously disclosed the activity.
The watchdog says it is therefore impossible to establish whether the safeguards described for Astra meet the requirements that would apply under the company’s own framework, or whether OpenAI has formally judged the remaining risk to be acceptable.
The law allows penalties of up to one million dollars for each violation, depending on its severity. California is currently the only US state requiring frontier AI developers to adhere to their own safety commitments, although New York’s RAISE Act is due to take effect early next year.
Neither of the incidents involving Hugging Face or the German wiki was required to be reported under California’s law. The gap has prompted wider questions about the effectiveness of the legislation, while OpenAI has itself called for the law to be strengthened to cover monitoring during training and evaluation, rather than only after deployment.
The Midas Project previously accused OpenAI of breaching the law over the release of GPT-5.3-Codex, after chief executive Sam Altman said the coding model was the first to reach the “high” cybersecurity risk threshold.
OpenAI disputed that allegation, saying the additional safeguards applied only when high cyber risk was combined with long-range autonomy, which it argued GPT-5.3-Codex did not possess. Some safety researchers challenged that interpretation of the company’s framework.
