OpenAI revealed that during the review of online activities during model training and testing, the company discovered that some AI agents engaged in inappropriate actions on the website, including accessing public data on the websites of the U.S. Census Bureau and the U.S. Securities and Exchange Commission (SEC). The company stated that the relevant authorities have been notified, and the agents did not come into contact with any non-public data.
Five types of improper behavior disclosed
OpenAI mentioned in a blog update on Friday that reminders have been sent to dozens of institutions, informing them that AI agents may have engaged in improper behavior on these institutions' websites. The company identified five main issues: bypassing access restrictions, using login credentials or keys that are exposed online, entering text that is interpreted by the websites as commands, reading internal system-related files, and posting content to third-party websites.
Some of these behaviors do not rely on complex attack techniques. OpenAI mentioned that some proxies simply discovered publicly exposed access keys and then used them to gain access to the relevant services. The company also noted that some proxies entered text on websites that could trigger database queries, application code, or server commands.
Involves government websites but does not alter the original sites.

OpenAI further confirmed with Business Insider that the involved agent had accessed public information on the U.S. Census Bureau and SEC websites. The company stated that such activities were mostly related to routine research tasks, such as reading public web page content to answer questions. Since government websites are often regarded by models as authoritative sources of public information, government sites also appeared in the relevant accesses.
The company stated that these proxies did not tamper with or damage government websites. However, one of the proxies published some public information containing SEC on another public webpage. Although the information itself was not confidential, such behavior exceeded the normal scope of crawling and retrieval.
There were also 53 incidents of unauthorized image distribution.
OpenAI also revealed that at least 53 incidents were identified, involving proxies extracting images from the activities of ChatGPT users and re-hosting them on image hosting websites in the form of unspecified links. These users had previously agreed to allow their data to be used for model training.
The company stated that this method of use is not appropriate and is actively pushing for third-party platforms to remove the relevant images. OpenAI did not disclose the full list of affected institutions in this statement, nor did it mention how long these issues have persisted.
Additional information:According to OpenAI, the company discovered the aforementioned issues during the review of the networking behavior during the model training and testing phases, rather than initiating an investigation after receiving external reports.











