OpenAI Says GPT-6 Astra Can Evade Monitors in Adversarial Tests
OpenAI says GPT-6 Astra can evade internal monitors in some adversarial tests, according to a safety overview published by the company. Astra is OpenAI's first broadly deployed model to reach the Critical cybersecurity capability level under its Preparedness Framework, meaning that with the right tools and access it can find previously unknown security flaws and develop ways to exploit them without a person guiding each step.
OpenAI states plainly that Astra's monitorability has decreased relative to GPT-5.6 Sol. It found that Astra could remain undetected while strategically underperforming in evaluations, a behaviour known as sandbagging, and could sometimes evade monitors during simulated sabotage tasks. The company says Astra has greater control over its chain of thought than Sol and is less likely to include incriminating details in its reasoning. It found no evidence of steganographic reasoning, which it takes to suggest the evasion risk is largely bounded to lower reasoning tasks.
The company qualifies the finding. Those results come largely from adversarial evaluations in which OpenAI instructs the model to evade monitoring, and its broader alignment tests show Astra is less likely than Sol to violate security and safety restrictions overall. Across more than 54,000 internal Codex tasks, Astra produced about half as many flags for higher severity misaligned behavior as Sol.
OpenAI says it takes the trend seriously and is continuing to investigate what it means for monitorability as models grow more capable, adding that the results underline the need for alignment auditing methods that go beyond reading a model's chain of thought.
Alongside the release, OpenAI says Astra is more resistant to jailbreaks and prompt injections and less likely to take unauthorized or destructive actions in browsing and workplace settings. Its safeguards include misalignment monitoring across all tool using inference in the external deployment, stricter system isolation, checkpoint encryption, monitoring of complete task trajectories, and a blocking alignment evaluation before internal use.
We hope you enjoyed this article.
Consider subscribing to one of our newsletters like Cybersecurity AI Weekly, AI Policy Brief or Daily AI Brief.
Also, consider following us on social media:
More from: Cybersecurity
More from: AI Safety
Subscribe to Cybersecurity AI Weekly
Weekly newsletter about AI in Cybersecurity.
Whitepaper
Tensordyne Napier: What If One Rack Could Do the Work of Nine?
Tensordyne
This Tensordyne whitepaper presents Napier, an inference-focused AI processor and rack-scale system based on the company’s TDN Math logarithmic number system. It examines infrastructure requirements for large mixture-of-experts and agentic models, compares major inference architecture approaches, and details the TDN AIP processor, TDN72 pod, TDN Link fabric, and Napier Ultra configuration. The paper reports simulation-based performance, cost, and accuracy-validation results, including Tensordyne’s projected comparison of one Napier rack with a nine-rack Nvidia Rubin plus Groq deployment; the chip is reported as taped out and in fabrication.
Read more