Abstract
- Anthropic Released on Monday, Claude Sonnet. The input price is $2 per million token, and the output price is $10 per million token – which is on par with Sonnet at $5, but only half of the price of Opus at $5.5.
- According to Anthropic, Sonnet scored 70.6% on Terminal-Bench 4.0, which is higher than Opus's 66.4%; an independent testing institution, Artificial Analysis, also reached a similar conclusion, with scores of 63.6% and 59.6% respectively.
- Artificial Analysis ranks it second only to Opus with a score of 5.5, but it indicates that its token consumption on each task is higher than that of any model it has tested.
Anthropic released Claude Sonnet 5.5 on Monday, which is an upgraded version of Sonnet 5 launched in June. Anthropic indicates that the running speed of this mid-range model is over 30% faster than its predecessor.
However, the most prominent feature of this model lies in its coding ability. In the Terminal-Bench 4.0 test – which assesses whether the AI agent can complete complex professional tasks through autonomous command input, with scoring based on the proportion of tasks completed – Sonnet achieved a score of 70.6%. Opus scored 66.4%, while Sonnet only managed 10.3%.
In simple terms, this cheaper model completed more tasks. An independent testing organization, Artificial Analysis, conducted its own version of tests and came to the same conclusion: for Sonnet at 5.5, it was 63.6%; for Opus at 5.5, it was 59.6%; and for OpenAI's GPT-6 and Astra, it was 59.1%.
Artificial Analysis stated on social media that Claude Sonnet 5.5 ( max ) has made significant progress on Terminal-Bench, and ranked among the top models in both Terminal-Bench 4.0 and Terminal-Bench-Science tests. The institution mentioned that in Terminal-Bench 4.0, it scored 64%, which is 50 percentage points higher than Claude Sonnet 5 ( max ), and also slightly higher than Opus 5.5 and GPT-6.
The performance also depends on the setting of “effort (level of effort)”. This adjustment will make the model spend more time thinking in exchange for better answers and higher costs. Anthropic indicates that under the settings of High effort, Sonnet can perform on par with GPT-6 Sol at FrontierCode, while the cost per individual task is about one-fifth of the latter.
In the GDPval-AA test – which uses a Elo system similar to chess rating points to score real professional jobs in 44 different professions – Sonnet scored 1844 with a 5.5, and Opus also scored 1846 with a 5.5, which can basically be considered a tie. GPT-6 and Sol scored 1487.
Competitors have also matched their prices. Last week, OpenAI reduced the price of GPT-6 and Sol to $2 per million inputs and $10 per million outputs; the mid-range model GPT-5.6 and Terra was priced at $2 per million inputs and $12 per million outputs. Anthropic has not released the benchmark test results for Terra.
The problem lies in
Sonnet 5.5 is a “high-yielding” model. Under the settings of max and effort, it generates an average of about 193,000 token per test task, which is the highest level recorded by Artificial Analysis. This is approximately 60% higher than that of Opus 5.5. Calculated in this way, the cost per task is 7.60 US dollars, which is about 50% higher than that of Sonnet 5. This does not align with the claim made by Anthropic that “up to 30% in costs can be saved”.
The savings mentioned by Anthropic come from using lower settings: under the default settings in Medium effort, the company claims that Sonnet can achieve better encoding results than the latter at less than one-tenth of the cost of the latter's best encoding performance. Artificial Analysis indicates that High effort are the most cost-effective settings. For everyday users, this means that by keeping the adjustment settings at a lower level, they can obtain nearly flagship-level encoding capabilities at a fraction of the price of a flagship model.
The table published by Anthropic consists of self-reported company data, while Artificial Analysis tested a pre-release version that contained a vulnerability. Anthropic expects that this issue will not have a significant impact on the results, or it may have merely underestimated the scores slightly. Anthropic also indicates that in complex tasks that require sustained judgment, Opus with a score of 5.5 is still significantly stronger.
Claude Haiku 5.5, designed for high-throughput, cost-sensitive applications, is expected to be launched in the coming weeks.











