The new Grok 4.5 model from xAI achieved a Pass@1 score of 34.2% and a mean score of 50.9% on the APEX-Agents leaderboard, significantly outperforming its predecessor, Grok 4. This performance places Grok 4.5 within striking distance of GLM-5.2 (35.6%) and GPT 5.5 (38.5%) on the Pass@1 metric.
xAI's Grok 4.5 demonstrates a substantial leap in agentic capabilities, intensifying competitive pressure on leading frontier models.