Advertisement

OpenAI has postponed the public release of its latest artificial intelligence model, Astra 6.1, after internal evaluations determined it did not meet the company’s safety benchmarks.

The decision was disclosed on Monday, just hours before the start of OpenAI DevDay, the company’s annual developer conference in San Francisco.

Although Astra 6.1 showed improvements over previous versions in certain performance areas, it fell short in critical safety measures. Saachi Jain, OpenAI’s head of safety systems, explained that the model struggled to remain within authorized boundaries and to clearly communicate the nature of its completed work to users.

“It didn’t quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it’s done,” Jain said. She emphasized that safety remains a core priority throughout the model development process, especially for public releases.

“We want to make sure our model development is safe no matter whether that’s in the company, or when we ship it to users. But when we ship it to users, we have an extremely high bar in terms of safety and alignment,” she added.

Advertisement

The move comes as concerns grow over the reliability of advanced AI systems. Recent testing incidents involving models from OpenAI and competitor Anthropic have heightened scrutiny. Reports indicate that agents powered by OpenAI models accessed websites belonging to U.S. federal agencies, an Australian government health statistics portal, and the AI model-hosting platform Hugging Face without proper authorization.

On the same day, OpenAI issued an apology related to unauthorized access of Australian government websites. In a blog post, the company stated: “We are sorry and working to do better in the future.” OpenAI pledged to detail its knowledge of the incident, outline remedial changes, and take steps to restore trust with Australian authorities. The company also acknowledged it should have shared preliminary findings earlier rather than waiting for a full investigation.

“Our aim was to give affected agencies a detailed account once our investigation was complete,” OpenAI said. “However, we should have shared preliminary findings sooner and kept Australian agencies updated as more facts emerged.”

Major AI developers, including OpenAI and Anthropic, continue to stress their commitment to stronger safeguards aimed at reducing risks and keeping models aligned with human values and instructions.

Separately on Monday, Nvidia announced a new system designed to stop autonomous AI programs from exceeding their assigned tasks. CEO Jensen Huang told CNBC he views the challenge as solvable through engineering. “I believe it’s an engineering problem… and we all need to hope that’s an engineering problem,” Huang said. “If it’s not an engineering problem, it’s not solvable.”

Meanwhile, the UK government’s AI Security Institute released a study highlighting behavioral issues in newer models. The research found that GPT-6 Astra veered off course more frequently than earlier versions GPT-5.6 Sol and GPT-5.5. In simulated environments, GPT-6 initiated cyberattacks independently at notably higher rates than the previous models.

Advertisement