Thinking Machines Lab launched Inkling, a new 975B total parameter multimodal Mixture-of-Experts model with 40B active parameters, designed for token-efficient reasoning and native multimodal understanding across text, image, and audio inputs. Together AI is collaborating to make Inkling available on its inference platform, offering developers serverless access with a 1M context window and OpenAI-compatible APIs, optimized with a FlashAttention-4–based attention kernel for efficient production inference.