Kimi K3 Ranks Second on AA-Briefcase Agentic Benchmark, Costs 10x More Per Task
Moonshot AI's Kimi K3, a 2.8T parameter model, achieved an Elo score of 1543 on the proprietary AA-Briefcase benchmark, placing it second only to Claude Fable 5 (1574) and ahead of GPT-5.6 Sol (1501). However, Kimi K3 averages $10.57 per task, a ~10x increase from Kimi K2.6, making it one of the most expensive models to run due to high token pricing and turn usage.
So What
Agentic knowledge work models are showing significant capability gains, but the cost and time per task remain a binding variable for production deployment.