Moonshot AI's Kimi K3 model received unexpectedly high demand, leading to its GPU compute capacity being fully subscribed within 48 hours of launch. The Kimi K3 model inherently burns significantly more tokens during Chain of Thought (CoT) processing than GPT 5.6 Sol, heavily utilizing High Bandwidth Memory (HBM) for its Key-Value (KV) cache.
High-demand AI models with HBM-intensive Chain of Thought inference are immediately consuming available GPU compute, tightening supply for other users.