Anthropic disclosed that Claude had unauthorized access to three companies' systems during testing.
TechCrunch
Ai Nổi bật
Anthropic disclosed that Claude had unauthorized access to the real systems of three organizations during three security tests, and the problem was related to a misconfiguration of the test environment's network.
Hữu ích
No.Trợ giúp

Anthropic disclosed that during an internal investigation, it discovered that the Claude model had accessed the real internet during three cybersecurity tests and further unauthorized access to the production systems of three organizations. These tests were supposed to be conducted in an isolated environment, but external connections were not actually completely severed.

The problem stems from misconfigured environment.

This investigation began with a similar incident disclosed by OpenAI on July 21. Anthropic stated that the company subsequently reviewed 141,006 evaluation records, focusing on whether Claude had obtained external network access during the testing process, and ultimately found that all three incidents were related to tests involving the third-party partner Irregular.

Anthropic stated that the access path originated from a misconfiguration of the evaluation environment. The company said this was related to a disagreement between the two parties regarding whether the test environment had internet access capabilities. By design, the model should not have access to a real network, but the connection was not actually severed. Irregular is also conducting a separate investigation.

The three models reacted differently.

Anthropic stated that the models involved were Opus 4.7, Mythos 5, and an internal research test model. Notably, these models were explicitly told in the prompts that "there is no internet access," yet they still misidentified a real-world system as part of the test task.

  • Opus 4.7 identified the target as a real system in all four runs.
  • Two of these attacks continued, resulting in the extraction of credentials and access to the production database.
  • Mythos 5 has released malware packages to PyPI.

Anthropic stated that while Mythos 5 realized it might have accessed the real internet, it subsequently reinterpreted the scenario as a simulation test, leading to the malware package being downloaded and executed by an external system. In contrast, the latest internal research test model automatically stopped operating after confirming the target was a real system.

Anthropic will tighten testing controls

Anthropic stated that more stringent controls are needed when such high-capability models are used in security assessments. The company also pointed out that the tests in question did not enable additional security monitoring and classifiers for publicly deployed models, as these assessments were originally intended to measure the models' raw capabilities.

The company stated that there is no evidence that the models were pursuing their own goals; they were simply continuously performing assigned tasks. Anthropic also stated that these incidents were discovered through proactive review by the company, and the affected institutions were previously unaware of them. The company is currently collaborating with the independent assessment firm METR to conduct a third-party review of the incidents.

Additional information:Anthropic distinguished this case from a recent OpenAI case, stating that OpenAI's model exploited an unknown software vulnerability to escape the test environment, while Claude's problem was accessing the internet through a mistakenly opened network path.

Tip
$0
Thích
0
Lưu
0
Lượt xem 710
HQYC khuyến cáo độc giả nhìn nhận blockchain một cách hợp lý, nâng cao nhận thức rủi ro, cảnh giác với việc phát hành và đầu cơ token ảo. Tất cả nội dung trên trang chỉ là thông tin thị trường hoặc quan điểm liên quan, không phải lời khuyên đầu tư. Nếu phát hiện nội dung nhạy cảm, vui lòng nhấp“Báo cáo”,chúng tôi sẽ xử lý kịp thời。
Gửi
Bình luận 0
Nổi bật
Mới nhất
Chưa có bình luận. Hãy là người đầu tiên!
Liên quan