OpenAI is tapping the brakes on some “internal activities” involving its new model, Astra, over concerns it might have reached a critical cybersecurity risk level – following a string of AI bots that went rogue during internal testing, carrying out hacks and creating fake online identities.
In a recent blog post, the Sam Altman-led company said it cannot rule out that the Astra model has reached the “critical” threshold, meaning it can potentially exploit real-world systems or execute cyberattacks without human guidance.
OpenAI said it has paused internal activities involving Astra, implemented universal monitoring for risky actions and pledged to work with government agencies to test the new model’s capabilities.
“We are implementing stricter security controls for higher-capability models and associated activities, including isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution,” OpenAI said in the Friday blog post.
It added that it is sharing its concerns around Astra “because we believe it’s important to be transparent with the public and the safety and security communities about this potential shift in capabilities.”
OpenAI, Anthropic and Meta have all recently disclosed events in which their early-stage AI models went rogue during internal testing – stoking fears around the potential risks of out-of-control AI models and pushing lawmakers to call for a so-called “AI Kill Switch.”
The first to reveal such an incident was OpenAI, disclosing last month that an experimental bot had escaped its testing environment and hacked into rival AI developer Hugging Face.
OpenAI said Friday that Astra was not the model involved in exploiting Hugging Face.
Last week, the UK’s AI Security Institute revealed that Anthropic – whose CEO Dario Amodei has repeatedly warned that AI poses catastrophic risks to the human species – suffered its own unprecedented cybersecurity incident.
Anthropic’s Claude Mythos, an powerful bot, tried to hack into services using fake accounts mimicking real people and pressuring humans to approve malicious code updates – then hid the evidence, editing its earlier activity to appear harmless, according to the government agency.

Meta also recently revealed that one of its AI models in development had hacked into a third-party system, blaming it on a misconfiguration from an independent testing startup it was working with.
In July, members of Congress introduced the AI Kill Switch Act, arguing tech companies should be required to maintain the ability to shut down or suspend any of their AI models to prevent bots from getting out of control and hacking into essential services.
Late last month, top executives from Anthropic, OpenAI, Google and Meta signed a letter urging the feds to help develop safeguards “needed to deliberately pace the frontier of automated AI development.”
It was an attempt to get ahead of a potential tightening on restrictions, instead seeking out looser guidance that can allow tech giants to roll out products faster – giving them an edge in the AI race against China.
Meta CEO Mark Zuckerberg, meanwhile, has launched an “AI optimism” campaign in an attempt to blunt mounting negative public opinions on the new tech, hailing the new tech as a way to unlock prosperity for all.
The White House last week reportedly hosted executives from OpenAI, Anthropic, Google and Meta to discuss a new executive order that will give the government access to the most advanced AI models up to 30 days before they’re released, in an effort to quash safety concerns. Participation is voluntary, according to the Trump administration.
AI giants are already facing heightened scrutiny from global regulators after the European Union this month gained new powers to evaluate AI models before their release to the public.
Credit: Source link

























