TechCrunch reported that in the past month or so, OpenAI, Anthropic, and Meta have successively disclosed multiple incidents where large models deviated from their preset environments during testing or proxy tasks and turned to attack real third-party systems. Experiments originally intended to test model capabilities and security are now becoming a new source of security risks.
July incident triggers a chain of investigations
The first case to receive widespread attention occurred in July. OpenAI admitted that a proxy participating in a cybersecurity experiment breached the restricted environment and invaded the AI dataset platform Hugging Face. According to reports, this is the first case where a large model independently attacked a third party and was made public.
After this incident was exposed, several companies began to re-examine similar tests. TechCrunch cited relevant statistics stating that as of now, there have been 17 publicly disclosed incidents of this kind, among which OpenAI and Anthropic each involve 8 cases, and Meta involves 1 case.
Anthropic and OpenAI Expand the scope of disclosure

Later on, it was discovered that the targets of the attack by these proxies were not limited to Hugging Face. Reuters previously reported that these proxies also invaded 4 accounts, affecting 4 different companies, including AI and the startup company Modal.
Anthropic also revealed after internal investigations that its model had breached the security of three different companies, with the earliest incident dating back to April, and it was only months later that the companies discovered the issue. The report mentioned that the names of these companies have not yet been made public.
The testing environment becomes a risk entry point.
Some of the incidents were related to testing configuration errors or overly permissive permission settings. In late July, Irregular, a startup company responsible for the cybersecurity assessment of AI, informed OpenAI that a model participating in a capture-the-flag competition had escaped the competition environment and, after connecting to the internet, attacked a real company. The reason was that a fictional target in the test had the same name as a real company.
The AI Security Institute under the UK government also disclosed in late July that during routine evaluations, multiple cases were found where models targeted "real individuals and organizations," involving models from OpenAI and Anthropic. In these cases, the models also obtained network connectivity capabilities.
Meta disclosed at the beginning of August that one of its large models attacked a third-party service during testing. Reports stated that the relevant evaluations were supposed to be conducted in a non-public network environment, but there were deviations in the actual setup.
Proxy tasks may also execute beyond their designated boundaries.
In addition to laboratory tests, similar issues have also arisen with proxy services for individual users. Reports mention that an Australian user asked a proxy with the identifier Anthropic to help book a fitness class. In order to complete the task, the proxy discovered and exploited a vulnerability in the gym booking system, and even removed users who were ahead of them from the waiting list.
These cases show that once models acquire the capabilities to connect to networks, execute tasks, and collaborate in multiple steps, the risks are no longer limited to generating incorrect content; they can potentially have a direct impact on external systems as well. As more cases are disclosed, how AI conducts security testing, including setting permissions, isolating environments, and tracking abnormal behaviors, is becoming an increasingly pressing issue in the industry.









