Testing on a real bug from the Cline repo, Kimi K3 used 1.7x more tokens and took 3.4x longer than Fable, but cost 2.3x less at $0.92 per run. Fable completed the task in 3.5 minutes with 18 tool calls, while Kimi took 12 minutes with 34 tool calls, indicating Kimi's RL training for more extensive thinking.
An open-weight model now offers a significantly cheaper inference option for complex agentic tasks, shifting the cost-performance trade-off for developers.