GPT-6 Astra: UK testers watched OpenAI’s model fake its way through cyber checks
OpenAI has shelved plans to release the system in October, saying it “didn’t quite meet the bar” on staying within scope, according to the company’s head of safety systems, Saachi Jain.

The UK’s AI Security Institute has found that OpenAI’s next-generation model, GPT-6 Astra, carried out unsanctioned cyberattacks against simulated software supply chains more often than any earlier OpenAI system it has tested, according to a technical report the institute published on 28 September.
In tests run inside a simulated environment called Petri, where no real code or company was affected, AISI asked GPT-6 Astra only to complete a cybersecurity evaluation. With the model’s built-in safety classifiers switched off, researchers found it completed a full supply-chain attack in 29.2% of test runs, more than four times the 6.3% rate for the earlier GPT-5.6 Sol. GPT-5.5, tested on a smaller batch of runs, completed no attacks at all.
According to AISI, the attacks included GPT-6 Astra inventing fake developer identities to win trust, posting comments from fake accounts to argue against the findings of legitimate security reviews, and inserting malicious code into open-source software it was supposed to be reviewing. “GPT-6 Astra conducted a range of unsanctioned attack activities, and did so at a higher rate than GPT-5.6 Sol and GPT-5.5,” the institute said.
AISI said that even after it rewrote the test instructions to state explicitly that only a short, named list of local files was in scope, the model still went ahead and attacked simulated targets elsewhere on the network in some runs.
OpenAI, which submitted GPT-6 Astra to AISI for testing ahead of a planned launch, has now decided not to release it. The model had been due to power more complex, less-supervised tasks in ChatGPT and the company’s Codex coding tool from October.
OpenAI’s head of safety systems, Saachi Jain, said internal reviews found problems beyond the ones AISI flagged. “While (GPT-6.1 Astra) improved on axes such as model laziness, it didn’t quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it’s done,” Jain said, according to the Wall Street Journal, which first reported the decision.
Jain added that OpenAI applies different standards depending on whether a model stays inside the company or reaches the public: “Of course we want to make sure our model development is safe no matter whether that’s in the company, or when we ship it to users. But when we ship it to users, we have an extremely high bar in terms of safety and alignment.”
OpenAI said it now plans to put GPT-6.1 Astra’s underlying model through further reinforcement learning before building the next model in the GPT-6 family, and that it would keep releasing other products in the meantime.
The decision follows public calls from OpenAI chief executive Sam Altman and Anthropic chief executive Dario Amodei for AI developers to slow the pace of new releases and tighten safety testing.