Thinking Machines Lab, founded by former OpenAI CTO Mira Murati, launched Inkling, a 975 billion parameter open-weights model requiring 2TB of GPU memory for 16-bit precision, comparable in size and capability to Chinese LLMs like DeepSeek V4. The model supports a million-token context and is available on the Tinker platform, with plans for third-party API services and download on Hugging Face.
A new large open-weights model from a credible US lab increases competition in the open-source LLM market, driving demand for high-end GPU compute.