Anthropic Opus 5 ARC-AGI 3 Benchmark Win Questioned Due to Training Data
Anthropic's Opus 5 reportedly achieved a significant win on the ARC-AGI 3 benchmark, but this performance is attributed to specific training on environments resembling the puzzles. Human contractors were allegedly used to generate chain-of-thought data or update weights based on rewards, raising concerns about generalization.