DeepSeek-V4, available in Pro and Flash variants, achieves significant efficiency gains, with V4-Pro using 27% and V4-Flash using 10% of DeepSeek-V3.2's single-token inference FLOPs at 1M context. The redesign includes a CSA/HCA hybrid attention mechanism, Manifold-Constrained Hyper-Connections (mHC) for stable information flow, and Muon optimizer for training stability. The system-level changes span attention, KV cache, MoE communication, and post-training, making long-context inference cheaper.