Demand for Kimi K3's GPUs has reached current capacity limits over the past 48 hours, prompting the company to temporarily halt new subscriptions and prioritize compute resources for existing members. Moonshot recommends deploying Kimi K3 on supernode configurations with 64 or more accelerators for optimal inference efficiency.
Acute demand for AI inference compute is immediately outstripping available GPU capacity, forcing service providers to ration access.