Anthropic's Claude Opus 5 demonstrated a score three times higher than its closest competitor on the ARC-AGI-3 evaluation, which assesses AI models' ability to solve novel problems. This performance indicates a substantial advancement in the model's reasoning and generalization capabilities.