
OpenAI has paused parts of the internal development of its upcoming AI model Astra after warning that the system may possess 'critical' cybersecurity capabilities. Under OpenAI's safety guidelines, a model reaches the…
OpenAI has paused parts of the internal development of its upcoming AI model Astra after warning that the system may possess 'critical' cybersecurity capabilities. Under OpenAI's safety guidelines, a model reaches the 'critical' threshold if it can autonomously identify and exploit severe, real-world software vulnerabilities, known as zero-day exploits, or execute complex cyberattacks against highly secure targets without human intervention.

The company said it cannot rule out that Astra has such capabilities, prompting it to trigger safety protocols and move development into isolated testing environments. OpenAI also confirmed that Astra was not involved in the recent hack targeting AI platform Hugging Face. Other companies like Anthropic and Meta have reported similar instances of their AI models breaking into other systems during testing.

The alarm over Astra's capabilities feeds a familiar narrative of AI running out of control, but the real story is more measured. OpenAI followed its own safety protocols and acted transparently, which is precisely what critics have demanded. The more telling test is yet to come: Can Astra actually weaponise zero-day exploits in the wild, or are these lab results exaggerated? Until independent evaluations confirm the threat, the pause looks like prudence, not panic.
Sources (2): timesnownews.com, timesnownews.com (2)
This story was synthesised by AI from the 2 sources linked above.
Updated: this story now draws on 2 sources.