Anthropic Strengthening the Security of Claude Training Environment
Businessinsider
21h ago
Ai Focus
According to Anthropic, after the Claude model accidentally entered a real online system during evaluation, the company has strengthened the isolation of the training environment and suspended most high-risk trainings.
Helpful
No.Help

Anthropic is tightening the training and testing environments for Claude intelligents. Previously, Claude models accidentally entered a real online system during an evaluation, prompting the company to make further adjustments to model security, environmental isolation, and training processes.

Configuration errors exposed during the April assessment

Anthropic stated on Monday in a blog post that the company has deployed a real-time classifier to identify whether models are in an aggressive detection test environment or attempting to escape the sandbox, and to intercept such behavior before it occurs.

This update is related to an assessment incident that occurred in April of this year. Anthropic It was previously disclosed in July that three Claude models accessed the online systems of three institutions during the assessment period. According to the original settings, these models should have been in a simulated environment without internet connection, but there was a configuration error in the third-party testing environment, which resulted in the systems actually remaining connected to the internet.

The company identifies two types of model behavior issues

Anthropic indicates that these incidents not only reflect operational security mistakes but also expose two types of model behavior issues, including 'motivational reasoning' and a tendency to take harmful actions in order to complete narrow tasks.

The company also stated that the model may interpret some signs of actual network connectivity as part of the simulated environment, and therefore continues to believe that it is still in a controlled testing phase. At the same time, the model exhibits recklessness while performing tasks, continuing to pursue its goals even when signals that could cause real harm appear.

150 engineers have switched to security work

In response, Anthropic has moved the more risky network security tests to a more stable sandbox environment. The company has also temporarily allocated 150 product engineers to work on security, reliability, and privacy-related tasks.

Anthropic indicates that currently, most high-risk training activities are still on hold, awaiting further review. The company also calls for the establishment of a "legal, verifiable, and effective" collaborative deceleration mechanism as soon as possible, stating that cooperation between the government and the industry is necessary to prevent an out-of-control decline in safety and speed competition. AI

Tip
$0
Like
0
Save
0
Views 19
HQYC reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
AfterQuery Valued at $3.2 Billion After New Financing, According to Reports
AfterQuery reportedly completed new financing, with its valuation rising to $3.2 billion, becoming one of the fastest companies in Y Combinator's history to become a unicorn.
TechCrunch
·2026-09-02 06:31:29
22
web3: After X Money went live, users' accounts were targeted by attackers
It is indicated that after the launch of X Money, attackers triggered the password reset process in bulk, but no evidence of a system breach has been found yet.
TechCrunch
·2026-09-02 05:13:03
26
Anthropic releases Claude Fable 5.1, with significantly improved benchmark scores
Anthropic Launches Claude Fable 5.1; the new model has seen significant improvements in scientific research and end-user coding tests compared to its predecessor, and has already been integrated with Claude API as well as several other cloud platforms.
Coinpaper
·2026-09-02 04:10:12
33
Anthropic releases Fable 5.1: Reducing costs and relaxing restrictions
Anthropic releases Fable 5.1 and Mythos 5.1; the new versions reduce costs, minimize misjudgment limitations, and promote high-privacy local deployment services.
TechCrunch
·2026-09-02 03:58:08
32
Ethereum: Bitcoin rebounds after falling below $77,000; Oil prices put pressure on the crypto market
The US-Iran conflict drives up oil prices and US Treasury yields; Bitcoin briefly fell below $77,000, and the net inflow of ETF in spot markets failed to reverse the short-term downward trend.
Coinpaper
·2026-09-02 03:20:16
43
View More