SpaceXAI's Grok 4.5 Tops AutomationBench-AA, Outperforming Claude Fable 5 and Opus 4.8
Grok 4.5 scored 51% on AutomationBench-AA, completing 79.9% of task objectives while strictly passing 21.9% of tasks, the highest measured for both outcomes. It achieves this at $0.34 per task, roughly a quarter of the cost of Claude Fable 5 ($1.35) and Claude Opus 4.8 ($1.46). The model is highly token-efficient, using approximately 8,000 output tokens per task, less than a quarter of Claude Opus 4.8's 32,000, and resolves tasks in fewer turns with more parallel tool calls.