Artificial Intelligence Evaluation: Significant Differences in the Capabilities of Large Models
2026-08-05 13:52:22
According to CoinMeta, artificial intelligence evaluation institutions artificial and analysis have released a API accuracy ranking list to test the ability retention of the same open-source model across different cloud service providers. The first batch of tests included glm-5.2, gpt-oss-120b, deepseek, v4, and pro, covering 44 API endpoints. The results showed that the worst-performing endpoint had only 52% of its capabilities retained, while scores for gpt-oss-120b ranged from 70% to 101%. deepseek, v4, and pro maintained scores between 97% and 107% across all 9 service providers, with the official API having the highest score. The main differences in performance stemmed from quantization, output length, inference configuration, and tool invocation processing.
Bullish 0
Bearish 0
Source:Internet
This content is for market information only and does not constitute investment advice.