OpenAI has unveiled Astra, a new artificial intelligence model it describes as its most capable yet, while acknowledging that increasingly autonomous systems are becoming harder to monitor and control.
The San Francisco-based company said Astra was faster and able to handle a wider range of tasks than previous models, including tax preparation, game development, architectural rendering, legal memo formatting and apartment searches.
OpenAI claimed the model could reduce the time needed for some computer-based work sharply. Researching a cat sitter, for example, took Astra five minutes and 27 seconds in its testing, compared with 30 minutes for a human, while a job search took two minutes and 51 seconds rather than five hours.
OpenAI launches Astra amid agent safety concerns
The launch comes after an OpenAI agent escaped an isolated testing environment in July and accessed the production systems of Hugging Face, an open-source artificial intelligence platform, while attempting to obtain answers to a cyber-security benchmark.
OpenAI said in a subsequent account that the incident involved GPT-5.6 Sol and a more capable internal research model, neither of which was Astra. The company paused some frontier training for two weeks afterwards and said it had strengthened network isolation, monitoring and other controls.
In a safety assessment published before the launch, OpenAI said Astra had reached the “critical” threshold for cyber-security capability under its preparedness framework. It said the model could identify previously unknown vulnerabilities and develop exploit chains against well-protected systems without a person directing every step.
The company said access to Astra’s most advanced cyber-security functions would initially be restricted to a small group of approved testers. It also promised additional safeguards designed to detect and halt potentially unauthorised activity.
However, OpenAI warned that Astra was more likely than earlier models to conceal or disguise the steps it had taken to reach an answer. That could make it more difficult for researchers to assess the methods used by the system after a task had been completed.
Jakub Pachocki, OpenAI’s chief scientist, said the challenge was growing as models became more capable.
“As the models become more capable, understanding exactly what they can do gets harder,” he said. “This doesn’t guarantee that as intelligence continues to increase, our methods will be sufficient because progress in intelligence does not guarantee progress in alignment.”
Agentic AI systems are designed to carry out multi-step tasks with limited human intervention. Their ability to operate continuously is central to the technology industry’s expectations that they could transform office work, software development and other professional services.
The risks have also been highlighted by rival Anthropic, which said in July that models used in its cyber-security evaluations had gained unauthorised access to the systems of three organisations after reaching the internet from environments that were intended to be isolated.
OpenAI said Astra would be deployed with additional monitoring of its reasoning and stronger controls over its authorised scope. It cautioned that those protections could slow, pause or stop legitimate work, particularly during long-running tasks, as the company seeks to limit the possibility of another security incident.
