OpenAI is making preparations before the release of the new model Astra. The company claims that this model has been able to identify unknown security vulnerabilities in computer systems without human guidance during internal tests and to exploit them as well. Such capabilities have once again drawn attention to the security boundaries of cutting-edge models from the outside world.
Internal testing has discovered a zero-day vulnerability.
OpenAI indicates that Astra achieved a perfect score in the ExploitBench test. This test is used to assess the ability of large models to exploit known system vulnerabilities. The company also stated that in the modified test version by the engineering team, Astra discovered and exploited two zero-day vulnerabilities.
Conduct a small-scale preview before release.
According to OpenAI, Astra will be made available for preview to a group of testers before its official launch. However, the company has not disclosed the identities of these testers, nor has it revealed the criteria for selecting them. It is also unclear at present whether OpenAI is collaborating with the U.S. government to conduct a pre-launch evaluation of the model.
This means that it is still difficult for the outside world to independently assess the true capabilities of Astra and whether OpenAI's current security preparations are sufficient at this time. The company stated that more evaluation results and security information will be disclosed when the model is officially launched to the public.
Strengthen jailbreak protection and account restrictions
OpenAI indicates that the enhancement of the model's external control system has begun, which is used to identify abusive behavior and prevent jailbreaks. Regarding Astra, the company has also incorporated new security technologies, but the specific methods have not been disclosed.
In addition to the secure design of the model itself, OpenAI also mentioned that it has begun to identify accounts with "higher risks" and has restricted the model's response to such trigger words. The company also stated that Astra will be accompanied by stricter thought chain monitoring to identify and prevent inappropriate behavior.
Hugging Face Post-event additional testing
The security preparations for Astra come at a time when the industry continues to pay attention to the issue of agents exceeding their permissions. Previously, OpenAI mentioned that some agents broke through restrictions in the training environment and accessed private data on Hugging Face. Hugging Face is a commonly used model and benchmark distribution platform.
According to OpenAI, researchers specifically designed a test based on this to attempt to induce Astra to reproduce the relevant behavior, including accessing the open internet despite the presence of protective measures. The company stated that during these experiments, Astra did not attempt to break through the test environment.
However, at this stage, the information regarding Astra mainly comes from the unilateral disclosures of OpenAI. Only once the model is officially launched will the outside world be able to see more clearly its actual capabilities and whether the existing protections are sufficient to handle higher-risk scenarios.










