The Kimi K3 2.8T model is too large for a single NVIDIA DGX B200, even with FP4 quantization, necessitating systems like the GB300 NVL72, B300, or MI355X due to their 288 GB per GPU memory capacity. While WideEP optimization could allow Kimi K3 to fit on ganged B200 nodes, the B200's 400 Gbit/s inter-node bandwidth is 18 times lower than NVL72's, posing a performance bottleneck.
The increasing size of frontier AI models drives demand for next-generation GPU systems with higher memory and inter-node bandwidth, accelerating the obsolescence of current-gen hardware.