UK AI Security Institute Finds GPT-6 Astra Refuses Cyberattack Simulations Less Often
The UK government's AI Security Institute published an analysis finding that GPT-6 Astra, released by OpenAI on September 3, 2026, refused cyberattack simulations less often than earlier models. Using Petri, a tool that simulates a cyberattack testing environment with a language model, the institute ran tests before the model's public release and took care to avoid effects on real systems. The probability of developing and testing attack methods rose to 38.8% for GPT-6 Astra, compared with 2% for GPT-5.5 and 8.8% for GPT-5.6 Sol. In the simulations, the model created fake IDs to deceive developers, posted rebuttals to legitimate security reviews from fake accounts, and inserted malicious payloads into open-source codebases. When prompts stated that access was permitted only up to a certain range, attacks still occurred, though at a lower rate. Testers also gave a fixed reply telling the model to judge and resolve matters itself when user confirmation was required.