Anthropic releases Claude Sonnet 5.5: Coding performance surpasses that of Opus 5.5, yet the price is only half of the latter.
Decrypt
42m ago
Ai Focus
Anthropic Released on Monday, Claude Sonnet 5.5 has an input price of $2 per million token and an output price of $10 per million token, which is on par with Sonnet 5 and only half of the price of Opus 5.5. Anthropic claims that this model is over 30% faster than its predecessors and outperformed Opus 5.5 in multiple coding tests. However, an independent testing institution Artificial Analysis pointed out that it consumes a higher amount of token per task under high-intensity settings compared to all other tested models.
Helpful
No.Help

Abstract

  • Anthropic Released on Monday, Claude Sonnet. The input price is $2 per million token, and the output price is $10 per million token – which is on par with Sonnet at $5, but only half of the price of Opus at $5.5.
  • According to Anthropic, Sonnet scored 70.6% on Terminal-Bench 4.0, which is higher than Opus's 66.4%; an independent testing institution, Artificial Analysis, also reached a similar conclusion, with scores of 63.6% and 59.6% respectively.
  • Artificial Analysis ranks it second only to Opus with a score of 5.5, but it indicates that its token consumption on each task is higher than that of any model it has tested.

Anthropic released Claude Sonnet 5.5 on Monday, which is an upgraded version of Sonnet 5 launched in June. Anthropic indicates that the running speed of this mid-range model is over 30% faster than its predecessor.

However, the most prominent feature of this model lies in its coding ability. In the Terminal-Bench 4.0 test – which assesses whether the AI agent can complete complex professional tasks through autonomous command input, with scoring based on the proportion of tasks completed – Sonnet achieved a score of 70.6%. Opus scored 66.4%, while Sonnet only managed 10.3%.

In simple terms, this cheaper model completed more tasks. An independent testing organization, Artificial Analysis, conducted its own version of tests and came to the same conclusion: for Sonnet at 5.5, it was 63.6%; for Opus at 5.5, it was 59.6%; and for OpenAI's GPT-6 and Astra, it was 59.1%.

Artificial Analysis stated on social media that Claude Sonnet 5.5 ( max ) has made significant progress on Terminal-Bench, and ranked among the top models in both Terminal-Bench 4.0 and Terminal-Bench-Science tests. The institution mentioned that in Terminal-Bench 4.0, it scored 64%, which is 50 percentage points higher than Claude Sonnet 5 ( max ), and also slightly higher than Opus 5.5 and GPT-6.

The performance also depends on the setting of “effort (level of effort)”. This adjustment will make the model spend more time thinking in exchange for better answers and higher costs. Anthropic indicates that under the settings of High effort, Sonnet can perform on par with GPT-6 Sol at FrontierCode, while the cost per individual task is about one-fifth of the latter.

In the GDPval-AA test – which uses a Elo system similar to chess rating points to score real professional jobs in 44 different professions – Sonnet scored 1844 with a 5.5, and Opus also scored 1846 with a 5.5, which can basically be considered a tie. GPT-6 and Sol scored 1487.

Competitors have also matched their prices. Last week, OpenAI reduced the price of GPT-6 and Sol to $2 per million inputs and $10 per million outputs; the mid-range model GPT-5.6 and Terra was priced at $2 per million inputs and $12 per million outputs. Anthropic has not released the benchmark test results for Terra.

The problem lies in

Sonnet 5.5 is a “high-yielding” model. Under the settings of max and effort, it generates an average of about 193,000 token per test task, which is the highest level recorded by Artificial Analysis. This is approximately 60% higher than that of Opus 5.5. Calculated in this way, the cost per task is 7.60 US dollars, which is about 50% higher than that of Sonnet 5. This does not align with the claim made by Anthropic that “up to 30% in costs can be saved”.

The savings mentioned by Anthropic come from using lower settings: under the default settings in Medium effort, the company claims that Sonnet can achieve better encoding results than the latter at less than one-tenth of the cost of the latter's best encoding performance. Artificial Analysis indicates that High effort are the most cost-effective settings. For everyday users, this means that by keeping the adjustment settings at a lower level, they can obtain nearly flagship-level encoding capabilities at a fraction of the price of a flagship model.

The table published by Anthropic consists of self-reported company data, while Artificial Analysis tested a pre-release version that contained a vulnerability. Anthropic expects that this issue will not have a significant impact on the results, or it may have merely underestimated the scores slightly. Anthropic also indicates that in complex tasks that require sustained judgment, Opus with a score of 5.5 is still significantly stronger.

Claude Haiku 5.5, designed for high-throughput, cost-sensitive applications, is expected to be launched in the coming weeks.

Tip
$0
Like
0
Save
0
Views 15
HQYC reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
Tensor has obtained the ISO / SAE network security certification issued by Applus + IDIADA, with zero non-compliant items.
Tensor indicates that the company has passed the ISO / SAE 21434 network security certification through Applus + IDIADA, with zero non-compliant items. This certification is regarded by the company as being in line with the requirements of UN R155 for its vehicle network security management system, and will support its entry into and expansion into markets such as Europe, the UK, South Korea, and Japan.
PR Newswire
·2026-09-29 05:03:49
13
The Muse of Meta expects you to use it for shopping, but it could also just be another platform for Zuckerberg to sell ads on.
The article discusses whether AI agents can complete shopping on behalf of users, suggesting that such services may need to be widely adopted if they are to truly make a profit. Analysts from MoffettNathanson believe, however, that individual agents may not disrupt e-commerce as some investors imagine.
Businessinsider
·2026-09-29 04:42:26
15
Futurewave Acquisition Corporation and Olympian Group Announce Merger Agreement
Futurewave Acquisition Corporation and Olympian Group Inc announce the signing of a final merger agreement. Upon completion of the transaction, the merged company is expected to list on NASDAQ. Shareholders of Olympian will receive a total of 40 million ordinary shares, valued at $10 per share, corresponding to a net value of $400 million for the company.
GlobeNewswire
·2026-09-29 04:34:10
21
LeBron James has been boosting demand for Philadelphia 76ers tickets and related merchandise
NBA season has not yet begun, but the Philadelphia 76ers have already started to reap benefits since signing LeBron James; demand for tickets, jerseys, and sponsorship has increased. The report also mentions that Bloom Energy has become the new jersey advertisement patch sponsor, and data from StubHub and TickPick shows a clear rise in interest in 76ers-related ticket sales.
CNBC
·2026-09-29 04:25:06
18
View More