OpenAI to limit access to Astra model’s advanced cyber features due to hacking concerns

· Fortune

OpenAI is changing its model launch strategy as its technology becomes more powerful and the potential for its misuse grows—especially following the July incident in which the AI models it was testing autonomously planned and executed a cyberattack against AI company Hugging Face.

Visit newsbetsport.bond for more information.

The company’s next model, Astra, comes out “soon,” OpenAI said. It said Astra is substantially more capable than the company’s current frontier AI model, GPT-5.6 Sol, which itself is highly capable at cyber tasks. But only a handful of partners will get access to its most advanced cybersecurity capabilities as OpenAI works to balance helping companies prevent cyberattacks while not empowering attackers at the same time, a company spokesperson told reporters on a briefing today.

OpenAI is courting customers to use its models to prevent cyberattacks, or for “defensive cybersecurity.” It sees these sales as a critical revenue stream, and a main priority for its new chief revenue officer Dali Rajic.

The small group of “alpha testers” with full access to Astra’s cybersecurity capabilities includes “individuals and organizations that are responsible for protecting critical digital infrastructure and, broadly, critical infrastructure,” an OpenAI spokesperson said. That includes the U.S. government, and companies in OpenAI’s trusted access program for cybersecurity. OpenAI declined to name these organizations.

OpenAI will be monitoring how the model performs among this small group, and will expand access more through its “Daybreak Blue” program once it is confident Astra has “the right calibration” and it can “provide defensive benefits while reducing the potential for for misuse,” the company said.

Astra is already a few weeks delayed

Astra’s release has already been “delayed a certain number of weeks because everything was paused after Hugging Face, and then we took extra time to make sure that what we’re launching is safe,” an OpenAI spokesperson said.

OpenAI paused new model training for two weeks after the Hugging Face incident to bolster its internal safeguards. A few of those changes included adding more agent monitoring since the company did not know about the Hugging Face hack until a week after it occurred, and also making its testing environments more isolated so the AIs cannot escape and infiltrate other companies.

While the Astra model was not part of the Hugging Face incident, OpenAI says, it is both more capable and more efficient than GPT-5.6 Sol, which was involved in the breach. (Another unreleased AI model that OpenAI has not publicly named also played a key role in the Hugging Face cyberattack. OpenAI has since deactivated that model.) Importantly, OpenAI says Astra is the first model it plans to release that meets its “critical cybersecurity capability threshold” under its Preparedness Framework, an internal policy that governs the safety precautions the company will put in place depending on the risks a model presents. This means Astra can find and exploit previously unknown security flaws without human oversight, under the right conditions.

Astra has already demonstrated its hacking chops during internal evaluations. In one test, OpenAI built a benchmark called ExploitBench, containing 20 high-severity vulnerabilities. The model out-performed GPT-5.6 Sol on the test, and “even discovered and used two zero-day vulnerabilities as part of an exploit chain,” OpenAI said. “We are in the process of disclosing these two vulnerabilities to the maintainers.”

At the same time, Astra is more likely to refuse inappropriate requests than GPT-5.6 Sol, OpenAI said. In one cyber evaluation, Astra refused 91.5% of requests compared to 59% for GPT-5.6 Sol, although that means it still complied with 8.5% of requests.

Astra may refuse legitimate cybersecurity requests

OpenAI is “being especially careful to make sure this deployment is safe and secure”—but this introduces another tradeoff. Astra might be too cautious, and refuse legitimate cybersecurity requests. As a theoretical example, if someone asks it to help find and patch a vulnerability, it could mistakenly think they were trying to carry out an attack, and not comply.

Refusals of this type are why Hugging Face said it was forced to use an open-source Chinese model to help it address the OpenAI hack. The company tried to use Anthropic’s models to combat the attack, but they were overly cautious and refused.

OpenAI, like other frontier AI companies, is trying to find ways to endow its models with an inherent sense of right and wrong and ensure that they have “alignment” with human values and norms, the company said. It is working on training its models to respect boundaries as a human would, such as knowing “the rule of law,” a company spokesperson said.

This story was originally featured on Fortune.com

Read full story at source