Qwen3.8-Max releases the official version, with a score only slightly higher than that of DeepSeek-V4-Flash.
2026-08-06 18:18:37
According to CoinMeta, Alibaba has released the official version of Qwen3.8-Max, which is a multimodal flagship model with 2.4T parameters, focusing on programming agent and professional office tasks. API has already been launched. In the four tests of terminalbench 2.1, deepswe 1.1, nl2repo, and toolathlon verified, Qwen3.8-Max scored only 1.7 to 3.9 points higher than DeepSeek-V4-Flash. After retesting, Qwen3.8-Max's score increased from 53 to 56 points, tying with Opus at 4.8, but it is still 1 point lower than Kimi K3. Qianwen scored 1739 points on the agent task, surpassing Kimi K3's 1685 points and GPT-5.6 SOL's 1730 points, trailing only by 5 points behind Claude Opus. Although the unit price of token has decreased, the average cost of completing the comprehensive evaluation task has risen from $0.53 to $1.14, which is higher than Kimi K3's $0.86. The reliability of knowledge has decreased, and the illusion rate has risen from 23% to 40%.
Bullish 1
Bearish 0
Source:Internet
This content is for market information only and does not constitute investment advice.