Emad Mostaque, co-founder of Stability AI, stated that the current high inference cost of the Chinese Kimi K3 model is due to immature infrastructure, not technical limits, and will not last long. He expects US-based specialized infrastructure companies like Fireworks, Modal, and Baseten to optimize Kimi K3's kernels, routing, quantization, batching, memory use, and serving systems. Kimi K3 currently uses twice the tokens for the same task compared to GPT-5.6, but these optimization efforts are expected to close that efficiency gap.
Inference cost for new models can drop dramatically with optimization, shifting the competitive landscape for open-source models and inference providers.