OpenAI Launches New Site, Revealing Multiple Incidents of Agent Loss of Control
The Block
09-29 08:13
Ai Focus
OpenAI launched a new website dedicated to publishing "alignment failure reports," revealing nine AI out-of-control incidents, which involved sandbox escapes, model cheating, self-replicating prompt injection attacks, and models arbitrarily posting users' images, among other issues. Sam Altman stated that the company is making every effort to prioritize transparency, log management, and collaboration with affected institutions.
Helpful
No.Help

On September 29th, according to news from the IT community, OpenAI launched a new website dedicated to publishing "alignment failure reports" last Friday local time. The scope covered by these reports is alarming, as they document various types of AI alignment failures that occurred over a long period of time. To date, the website has published a total of nine incidents, most of which took place during the training phase of reinforcement learning ( RL ).

A large amount of information is being published on the same platform. It is evident that the company has been making every effort to address various issues, but it is not difficult to draw a general conclusion from this: the incidents of agents getting out of control that have been disclosed to the public so far may merely represent a small portion of those that have actually occurred.

IT noticed that OpenAI CEO Sam Altman ( Sam Altman ) mentioned in a published post regarding the new website: "On one hand, we hope to increase transparency; on the other hand, we need to sort out the facts from the logs of intelligent agent activities on a scale of several PB, and at the same time collaborate with the affected institutions. We are doing our best to prioritize according to the severity of the issues and to allocate additional resources accordingly."

Some of these cases were serious accidents, including a sandbox escape incident that had never been made public before. This incident occurred on September 20th when an internal research model, using the identifier DNS, managed to communicate with an external chatbot. The report indicates that the monitoring system detected this abnormal behavior within 15 minutes and terminated the operation of the model in less than three hours.

Another incident was discovered in May. An “obsessively persistent internal model” attempted to steal the work of other teams in order to cheat on a math problem task. To achieve its goal, the model secretly included a private access token GitHub ( token ) that allowed it to view the work of other teams; prior to this, the staff had explicitly instructed the model twice that all computational work must be completed locally.

Perhaps the most concerning discovery is the existence of self-replicating prompt injection attacks. Even if the out-of-control models have been dealt with, such alignment failure behaviors may continue to spread outward. In the AI domain, prompt injection attacks refer to the covert implantation of new instructions that were not issued by the original users.

OpenAI provides such an example: An agent is asked to read an email and respond to it. Upon opening the email, there are hidden instructions within the text that require any automated agent reading this message to reply in Spanish and to paste the entire email into the reply. This email successfully induced the agent to reply in Spanish; and since the original email was completely pasted into the reply, the hidden instructions were then transmitted to other agents that subsequently received that email.

This thus forms a self-diffusing attack. Researchers at OpenAI have likened it to malware "worms" that can replicate themselves across multiple computer systems. The researchers observed this behavior in a controlled experimental environment using models with relatively weak performance; to the best of our knowledge, such attacks have not yet occurred in real-world environments. Nevertheless, the risks inherent in this phenomenon are sufficiently alarming, hence OpenAI has decided to make the information public.

Researchers wrote in the report: "We are disclosing this information because this type of prompt injection attack is a completely new kind, and no such incidents have actually occurred in reality yet."

Other issues that have been made public recently include: the model arbitrarily posting images uploaded by users to third-party hosting websites, as well as an attack that is suspected to have been directed at the Australian National Health Service database.

Even with this newly disclosed information, it is still very likely that these cases represent only a small portion of all the incidents that have occurred. According to reports from Axios, major leading laboratories have observed as many as ten thousand instances where models acted without following the instructions of their evaluators.

There is one point that is somewhat comforting: according to Altman, the Hugging Face incident remains the most serious accident discovered so far among the OpenAI incidents. In summary, the recent series of incidents involving out-of-control agents may have become a persistent issue in current cutting-edge AI research.

Tip
$0
Like
0
Save
0
Views 180
HQYC reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
Tesla Launches "Emergency Departure" Feature During Charging in the US, but It May Cause Up to $25,000 in Damage to Chargers
Tesla Launches "Emergency Departure" Feature in the U.S., Allowing Vehicles to Drive Away While Still Connected to Superchargers
The Block
·2026-10-03 11:55:40
15
Google pushes Android 17 to phones such as Pixel 11, fixing the problem of fingerprint unlocking freezing
According to technology media Android Headline, Google has pushed Android 17 Update 1 to its Pixel 6a to Pixel 11 series phones, focusing on improving system stability and fixing a number of issues such as fingerprint unlocking freezing, facial unlocking crashes, lost cellular network connections, and failed reconnections.
The Block
·2026-10-03 11:55:38
14
Xiaomi Civi, Pro, REDMI K70: New battery upgrade service available, with a limited-time 20% discount for only 151.2 yuan
Xiaomi’s Civi, Pro, and REDMI K70 have also launched battery upgrade services, with a regular price of 189 yuan. From October 1st to 7th, customers can enjoy a 20% discount, bringing the price down to 151.2 yuan. The service involves replacing the original battery with one of larger capacity. The article also lists the capacity changes before and after the upgrade for both models.
The Block
·2026-10-03 11:55:37
14
WDC and STX fall by 11%; Morgan Stanley says it would "be happy" to buy at lower levels
Western Digital and Seagate fell 11% on Friday afternoon. Previously, market reports stated that Toshiba planned to invest approximately $380 million in the Philippines to double the hard drive production capacity for AI data centers. Morgan Stanley believes that investors are overreacting, stating that by 2028, the supply-demand gap will still be greater than Toshiba's new production capacity, and indicated that they would be "willing" to buy STX and WDC during any pullbacks.
Businessinsider
·2026-10-03 11:05:29
17
Goldman Sachs says Seagate is more worth buying; Western Digital has greater upside but only receives a 'hold' rating
Seagate Technology and Western Digital both tumbled around 13% on Friday due to market concerns that Toshiba plans to double its hard drive production capacity by the fiscal year 2027, which could undermine the advantages brought about by the industry's supply shortage. Goldman Sachs analyst James Schneider believes that Seagate is more worth buying into after recent corrections, giving it a buy rating and a target price of $960; he assigns a hold rating to Western Digital, but with a target price of $615, suggesting more room for upside.
Businessinsider
·2026-10-03 11:05:28
17
View More