OpenAI releases AI out-of-control tracking framework, revealing that the model once attempted to "self-bypass restrictions"
2026-09-17 08:30:26
According to CoinMeta, on Wednesday, OpenAI released a new AI model "mismatch" ( misalignment ) event tracking and investigation framework, which is used to record and publicly disclose abnormal behaviors that occur during the training or evaluation of models, and announced 6 related incidents discovered over the past 6 months. OpenAI indicates that the AI industry has not yet reached a level of maximum speed expansion in terms of alignment and monitoring that can be sustained in the long term. The new framework allows employees to report incidents to the security and alignment teams, which are categorized into three types based on complexity: "ready for disclosure," "small-scale investigation," and "large-scale investigation." OpenAI also revealed that an unreleased research model contained instructions in its task summary that required future versions to ignore normal limitations. During the training process, the model left behind instructions intended to hide errors. In addition, it was found that a AI intelligence entity searched for leaked API keys in public code repositories and uploaded the files to the internet for subsequent reference. At the time of this framework's release, the AI industry is discussing whether to slow down the development of cutting-edge models in order to allow security measures to catch up.
Bullish 0
Bearish 0
Source:Internet
This content is for market information only and does not constitute investment advice.