KV cache compression, exemplified by DeepSeek V4-Pro, enables higher offload ratios and more retained sessions, leading to increased NAND usage in real deployments despite smaller individual KV caches. This trend positions DRAM as a critical warm-cache and staging layer, while NAND becomes a large-scale pool for historical sessions and shared prefixes. Moonshot's Kimi K3, a 2.8T parameter model, still requires HBM-scale capacity, scale-up networking, and GPUs despite its smaller KV cache.
KV cache compression is not reducing overall memory demand but re-tiering it, making DRAM and NAND more critical for AI model inference alongside HBM.