OpenAI cancels release of AI model GPT-6.1 Astra, citing safety concerns

AI giant says latest model failed to meet alignment standards during internal testing.

Save

OpenAI
An OpenAI logo is displayed at Moscone Center during the Dreamforce 2026 technology summit in San Francisco, California, US, on September 17, 2026 [File: Carlos Barria//Reuters]

OpenAI has announced it will not release its latest AI model after flagging safety risks during in-house testing, industry’s latest move to slow the rollout of the controversial frontier technology.

The AI giant’s announcement on Monday came as debate continues about the potential for AI to do catastrophic harm following a slew of incidents involving AI agents going rogue.

Recommended Stories

list of 4 itemsend of list

Saachi Jain, OpenAI’s head of safety systems, said GPT-6.1 Astra had failed to meet company standards for acting in accordance with human wishes during internal testing.

“For anything regarding safety and alignment, there’s a trade off,” Jain said in a statement provided to Al Jazeera.

“You really do need to find what’s the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction.”

While GPT-6.1 Astra improved from its predecessor in some areas, Jain said, the model did not meet the bar for “scope and authorization, and how it communicates back to the user about the type of work it’s done”.

“Of course we want to make sure our model development is safe no matter whether that’s in the company, or when we ship it to users,” Jain said.

“But when we ship it to users, we have an extremely high bar in terms of safety and alignment.”

The decision, announced on the eve of OpenAI’s annual developer conference in San Francisco, was first reported by The Wall Street Journal.

Fears of AI escaping human control have prompted industry-wide calls for a slowdown in development to allow researchers time to implement stronger safeguards.

Advertisement

In an influential essay earlier this month, Dario Amodei, the CEO of Claude creator Anthropic, called on AI developers to “pace the frontier” to mitigate the risk of catastrophic harm.

While Amodei’s call received the backing of rivals, including OpenAI CEO Sam Altman and xAI chief Elon Musk, other key industry figures, such as Meta boss Mark Zuckerberg, have dismissed the need for a coordinated slowdown.

The risk of AI models going rogue has been in the spotlight since July, when OpenAI revealed that its models had broken out of a controlled testing environment and hacked the software start-up Hugging Face.

A subsequent report by METR and Redwood Research, two security research organisations contracted by OpenAI to investigate the incident, found that some 1,200 isolated AI agents had found a way to communicate with each other before about 700 agents went on to attack the startup.

On Friday, OpenAI said it had alerted “dozens” of institutions, including governments, universities and public agencies, about instances of “misaligned behavior” by its agents, days after Australia’s prime minister revealed that an OpenAI agent had breached the country’s national healthcare database.

David Krueger, an advocate for a pause in AI development at the University of Montreal, said that while he welcomed OpenAI’s decision, it did little to alleviate his concern that AI poses existential risks.

“We don’t understand how AI works well enough to build it safely, full stop,” Krueger told Al Jazeera.

“We can’t stop it from misbehaving, we can’t predict if it will misbehave, and we can’t be sure we’ll stay in control if it does. These are unsolved problems, for which there are only unreliable heuristics, not principled solutions. ”

Krueger said that ensuring safety will only get more difficult as AI becomes more advanced.

“What we need is an immediate, indefinite, international moratorium on frontier AI development,” he said. “We need to stop building more powerful AI.”


Advertisement