Anthropic's Claude Opus 5 model scored three times higher than its closest competitor on the ARC-AGI-3 evaluation, which assesses AI models' ability to solve novel problems.