The new Inkling model employs a sliding-window attention on 5 of every 6 layers and 8 KV heads, resulting in a roughly 6x smaller KV cache compared to all-global designs at 1M-token contexts. This architectural choice aims to enable faster and cheaper long-context inference by reducing memory bandwidth requirements during decoding.