The Kimi K3 model, with 2.8 trillion parameters, requires approximately 1.4TB of HBM even with 4-bit weights, necessitating a full rack of GPUs for operation. Despite efficiency gains like a smaller KV cache, the model's enlarged 'brain' ensures HBM capacity remains fully utilized.
Model scaling continues to absorb all available HBM capacity, making efficiency gains irrelevant to overall memory demand.