ARC-AGI-3 extracts two rankings: the naked model accounts for less than 1%, while with the addition of the Agent framework, it reaches 36% on the first day.
2026-07-30 15:40:03
According to CoinMeta, the paper ARC-AGI-3 divides the evaluation into two separate rankings. The score of the bare model was less than 1%, but with the addition of the Agent framework, it reached 36% on the first day. The official ranking prohibits external harness; all models use the same minimalist prompt words, and no tools are provided. What is measured is the model's native intelligence without the support of these frameworks. In contrast, the community ranking allows for the use of harness; scores are reported by the models themselves, and ARC Prize does not perform any verification by default. The paper explicitly warns that "scores from the community ranking should not be interpreted as evidence of AGI progress." On the official ranking, the scores of cutting-edge models are all below 1%. OpenAI claims that the evaluation framework hinders the performance of GPT-5.6, but Opus using the same framework scored 2.3 times higher. OpenAI then re-performed the evaluation using their own Responses API framework, retaining historical reasoning and compressing overly long contexts, resulting in a score increase from 13.3% to 38.3% for GPT-5.6. However, the issue is that Opus also used this same general framework and scored 30.2%. While the framework indeed hindered the performance of GPT-5.6, it did not have the same negative impact on Opus.
Bullish 0
Bearish 0
Source:Internet
This content is for market information only and does not constitute investment advice.