The Kimi Delta Attention (KDA) mechanism uses a fixed-size memory, or "fast weights," that retrain on every token during inference, allowing for efficient processing of long contexts. This innovation enables up to 6x faster and cheaper throughput at 1 million context while outperforming full attention on evaluations. Moonshot AI can now offer flat pricing beyond 200,000 context, unlike other models with tiered pricing.