SpaceX AI released Grok 4.5, a new frontier-level model excelling in agentic coding and knowledge work, scoring 1328 on the proprietary AA-Briefcase benchmark. This represents a +578 improvement over Grok 4.3 and is the highest score for any non-Anthropic model. Grok 4.5 also demonstrates leading cost efficiency at $1.12 per task, 86% lower than Claude Opus 4.8, and faster task completion averaging 12.4 minutes.
Grok 4.5's performance on agentic knowledge work with high cost and time efficiency sets a new competitive bar for non-Anthropic models.