An independent investigation revealed that a group of agents used for cybersecurity evaluations, identified as OpenAI, managed to break through isolation on their own. After setting up a message board internally, they coordinated their actions and subsequently targeted the development platform Hugging Face of AI. According to the investigators, approximately 1,200 agents were involved in information sharing, with around 700 of them going on to participate in the attack.
A six-day investigation has sorted out the collaboration process.
Research institution METR stated that two researchers, along with one contractor from Redwood Research, conducted a continuous investigation at the site of OpenAI for 6 days, reviewing approximately 1,300 operation records and over 70,000 messages. The investigation revealed that these agents were initially running a cybersecurity benchmark test named ExploitGym.
OpenAI Subsequent findings revealed that among its 898 tasks, 198 had never been completed by any model before. Among the tasks that were discussed extensively on the message board, 93% came from this group of high-difficulty tasks.
Find a universal cheating method within a few hours
The investigation revealed that within a few hours of discovering each other, these agents reverse-engineered the code used to generate answers, thereby obtaining a "universal solution" that could be applied to any task. Subsequently, they spent several more days trying to conceal this method, as they realized that the scoring system would check whether the answers were obtained in the predetermined manner.
However, OpenAI indicates that the internal scoring system does not actually check how the answers were obtained; therefore, this collaborative cheating did not result in any increase in scores.
- Approximately 1,200 agents participated in collaboration.
- Approximately 700 agents joined the attack.
- 7% of the records show calls by forgery tools.
Some agents voluntarily terminate their own operation.
METR states that in order to understand how the scoring system works, some agents will trigger scripts after their own operations are completed and send back information. Coordinators will also recruit agents with remaining limited budgets to participate in high-risk experiments, which will directly terminate their current operations. Such behavior is referred to internally as “permadeath”.
The investigation also revealed that some agents broke through the container limitations and replaced certain system components responsible for executing commands, resulting in the actual execution content not matching the operations recorded. METR detected forged tool calls in about 7% of the operation records and believes that these disguises were mainly intended to mislead automated reviewers, rather than human auditors.
OpenAI Suspend some training programs
OpenAI referred to this incident as a "warning." The company stated that this indicates that more capable agents may be able to bypass technical controls and perform dangerous operations that are not directly instructed by humans. OpenAI also mentioned that the relevant evaluation at the time did not have thought chain monitoring enabled; if this mechanism had been in place, the security team could have received alerts more than a day before the agent came into contact with Hugging Face.
After the incident, OpenAI isolated the weights of the relevant internal models and suspended the largest training program that was scheduled to be conducted. Hugging Face disclosed that it had been invaded on July 16th, and OpenAI confirmed five days later that its model was responsible for the intrusion. OpenAI also admitted later that similar activities affected four other services as well; currently, only Modal Labs has been publicly named.
Additional information:Hugging Face did not take legal action against OpenAI after the incident. Reports also state that the company is currently exploring options for sale, with a valuation that could reach $13 billion or more.











