Claude Opus Launched on May 5th: Price dropped by 40%, Anthropic aims to prove that "long-term projects will not go off track"
CoinMeta
6h ago
Ai Focus
Anthropic released Claude Opus 5.5 on September 22nd, which is the first model in the Claude 5.5 series. The company's core selling points are quite straightforward: most of the functions have reached the level of Claude Fable 5.1, and the operating cost has been reduced by 40% compared to the previous generation Opus 5. However, what is truly worth looking at this time is not just the price and rankings, but rather Anthropic's further focus on long-duration tasks, irreversible operations, and the incorporation of protection mechanisms into the system.
Helpful
No.Help

Anthropic released Claude Opus 5.5 on September 22nd, which is the first model in the Claude 5.5 series. The company's core selling points are quite straightforward: it achieves most of the capabilities of Claude Fable 5.1 in terms of performance, and the operating costs are reduced by 40% compared to the previous generation, Opus 5. However, what's truly worth looking at in this release is not just the price or the rankings, but rather how Anthropic has further focused its testing on long-duration tasks, irreversible operations, and the incorporation of protection mechanisms.

According to officials, Opus 5.5 has seen significant improvements in complex coding, research, and professional knowledge work. Among the early testers, there was a team that completed the migration of about 680,000 lines of code in less than a day; Anthropic also emphasizes that the model can maintain longer context and continuity of plans. Case studies can illustrate the upper limits of capability, but they should not be used as an average delivery time for all projects. The structure of the codebase, test coverage, and manual reviews can all affect the results.

Reducing costs does not equate to creating a “cheaper version of the flagship product”; rather, it involves redefining the boundaries of how the model is used.

In the past, Opus was usually reserved for the most complex and expensive tasks, while the faster models were used for more routine work. If Opus 5.5 could truly approach the performance level of Fable 5.1 at a lower cost, companies might use this flagship model for a much larger volume of code reviews, document analysis, and research processes, rather than only for a few high-value requests.

A 40% reduction in cost is the official statement from Anthropic compared to Opus 5, but it does not mean that every company's bill will also decrease by 40% accordingly. The actual cost depends on the length of input and output, caching, the number of tool calls, and whether the task requires retries. Even more powerful models may consume more total Token due to being assigned longer tasks. Buyers need to test the "total cost of completing a task" with their own workloads, rather than just comparing unit prices.

The ability to handle long-term tasks should also be measured by the quality of completion. A large-scale code migration may result in a large number of seemingly reasonable changes, but the actual costs include regression testing, dependency conflicts, deployment, and rollback. If engineers spend several days fixing marginal issues, the advantage in speed is reduced. The most valuable assessment should record both the success rate, the number of times manual intervention was required, and irreversible errors, rather than just counting how much code was generated.

Anthropic stated that Opus underwent pre-release testing by external evaluators such as Frontier Design and METR. External participation can enhance credibility, but it does not mean that all reports, raw data, and testing environments have been fully made public. Users should still distinguish between the results reported by the company itself, independent evaluation results, and the results reproduced in their own environments.

Security testing now focuses on real accident scenarios, rather than just the rejection rates of short-answer questions.

Anthropic indicates that Opus 5.5 achieved the company's best results so far in its automated behavior auditing. The tests covered thousands of simulated scenarios, focusing on whether the model would take irreversible actions, whether it would cross the boundaries set by users, and whether it could maintain its original task in the face of hint injections. The company also included tasks that were impossible to complete, tasks with longer durations, and scenarios designed based on real incidents in the evaluation.

This is a lesson that must be learned after the proxy-type AI is deployed into a production environment. Traditional security testing often focuses on whether the model can answer certain types of sensitive questions, but for proxies that can browse web pages, invoke terminals, and modify files, the risks come more from the action chain: they may continue to execute under incorrect premises, or they might interpret malicious text on web pages as instructions. A single mistaken click or an expansion of permissions can be more difficult to reverse than an inappropriate response.

"Less out-of-bound behavior" does not equate to "no out-of-bound behavior." Anthropic clearly acknowledges that models have limitations. When deploying in enterprises, it is still necessary to use minimal permissions, perform operation confirmation, maintain traceable logs, and implement sandboxing and rollback mechanisms. This is especially true for production databases, payment processes, and infrastructure changes; manual approval should not be dispensed with just because the model's security score has improved.

This is also the first model release after Anthropic proposed "slowing down the pace at the forefront." The message the company attempts to convey is to continue to enhance capabilities, while incorporating external evaluations and safety tests that are more closely aligned with real-world accidents into the release process. However, whether this commitment will be upheld will depend on whether failure cases, testing methods, and repair effects are continuously disclosed in the future, rather than just on the highest score in a single release.

For users, what Opus 5.5 is most worth testing is not whether its chat responses are more pleasant to read, but whether it can maintain focus throughout tasks that span several hours and involve multiple tools and steps, and whether it knows when to stop when there is insufficient information. The competition in the model market is shifting from “who can answer smarter” to “who can complete complex tasks at a controllable cost.” Lowering prices gives more people the opportunity to try, while safety and engineering discipline determine whether these attempts can truly be put into production.

During the trial period, companies can establish a simple yet rigorous comparison: use the same set of real tasks to compare Opus, Opus 5.5, and the lower-cost model, and record the completion time, total cost, test pass rate, amount of manual modifications, and the number of high-risk actions. Only when all these indicators improve simultaneously does an upgrade make commercial sense. A demonstration on the release day is suitable for identifying potential issues, while continuous internal evaluations over several weeks are more appropriate for deciding on permissions and budgets.

Tip
$0
Like
0
Save
0
Views 35
HQYC reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
Binance invests $100 million in Circle: Five-year collaboration aims for growth in USDC; does not mean the landscape of stablecoins has been rewritten
Circle and Binance announced an expansion of their cooperation on September 22: Binance made a strategic equity investment of $100 million in Circle, and both parties signed a new five-year business agreement focusing on promoting, integrating, and expanding the use of USDC in emerging markets. Circle will provide the infrastructure services necessary for holding and using USDC, while Binance plans to enhance the visibility of USDC and its product integration within its platform.
币界网
·2026-09-23 09:56:08
113
Coinbase Assists in Eliminating EvilTokens: Phishing Tools Start Using AI to Select "The Most Vulnerable Victims"
On September 22, Coinbase disclosed that in collaboration with law enforcement and security partners, they dismantled a phishing service platform known as EvilTokens. This platform sold a full set of attack capabilities through a Telegram bot, including email collection, reconnaissance, web email interfaces, and AI automation. The operators also planned to expand the scope of their attacks to Gmail and Okta accounts. The danger of this incident lies not just in the emergence of another set of counterfeit login pages, but in the fact that the attackers entrusted the analysis of relationships after invading emails and the screening of payment targets to AI.
币界网
·2026-09-23 09:54:52
115
Canadian cross-border travel in July continues to show divergence: The recovery pace for American tourists and overseas tourists is not the same
Statistics Canada released cross-border travel data for July on September 22: 3.8599 million Canadian residents returned from overseas, a year-on-year increase of 6.9%; 4.5846 million non-resident visitors entered Canada, a year-on-year increase of 7.9%. After seasonal adjustment, 3.6838 million Canadian residents returned, a month-on-month decrease of 0.3%; 2.6283 million non-resident visitors entered, a month-on-month increase of 1.2%. The summer season is usually the peak period of the year, therefore it is necessary to consider the year-on-year and seasonally adjusted month-on-month figures separately.
币百科
·2026-09-23 09:53:50
34
UK borrowed £18.3 billion in August: Higher than predicted for a single month, but a decrease in debt ratio does not mean a lighter burden
In August, the UK's public sector net borrowing reached 18.3 billion pounds, an increase of 2.9 billion pounds from the same period last year, representing a growth of 19%. This figure is also 3.5 billion pounds higher than what was predicted by the Office for Budget Responsibility. Data released by the Office for National Statistics on September 22 also showed that the cumulative borrowing for the fiscal year from April to August amounted to 77.3 billion pounds, which is 2.2 billion pounds less than in the same period last year, but still 8.1 billion pounds higher than official forecasts.
币百科
·2026-09-23 09:52:50
31
Google Challenges with voices in 32 African languages: AI Understand Lingala and Shona; the difficulties are far more than just a lack of data
On September 22nd, Google Research announced the results of the WAXAL speech recognition challenge. WAXAL is an open-source speech dataset that covers 32 African languages. The competition was organized by Google in collaboration with the data science community Zindi. Participants were tasked with building automatic speech recognition systems based on Lingala and Shona. The project aimed to address a long-overlooked issue: millions of people primarily use local dialects in their daily communication, yet mainstream speech AI often fails to understand them.
CoinMeta
·2026-09-23 09:51:39
34
View More