Anthropic Claude Fable 5 Achieves New Best Score of 91.9% on WeirdML Benchmark
Claude Fable 5 (max) consistently scores within a few percentage points of the SOTA on individual tasks, achieving SOTA on 7 of 17 tasks with its worst run only 7% behind the SOTA. The model completed these runs with only two attempts per task, demonstrating high consistency.
So What
Anthropic's Fable 5 model demonstrates strong, consistent performance on a new benchmark, indicating a capability step-up for its next-generation models.