The analysis refutes claims that Kimi K3 significantly reduces memory and networking demand, noting its large MoE architecture with 2.8 trillion parameters and 16 of 896 experts activated per token, requiring substantial HBM, scale-up networking, and GPUs. While Kimi K3 achieves 75% KV cache compression (DeepSeek V4 reached 90%), its total parameter count of 2.8T translates to approximately 1.4TB of weights, which does not fully occupy the 13.4TB HBM of a GB200 NVL72 system.
Model architecture choices like Kimi K3's MoE and KV cache compression shift memory demand between HBM and NAND, but overall compute and HBM needs remain high.