LMSYS Org and Thinky Machines released Inkling, their first open model, featuring 975B total parameters and up to 1M context, with native reasoning over text, images, and audio. The model incorporates deeply optimized architecture including ShortConv, a fused all-reduce kernel that is 2-3.6x faster, and full CUDA graph prefill for a 14-17% speedup. It also utilizes MXFP8 KV cache, doubling capacity while maintaining performance within 14% of bf16, with day-zero support on SGLang and Miles.