The Inkling model, with 975 billion total parameters (41 billion active), supports text, image, and audio input, ranking second among open-weights models on AA-WER at 3.5% behind Mistral's 24B Voxtral Small. It processes audio at approximately 11x real-time and is available via Thinking Machines' Tinker platform API for $6.60 per 1,000 minutes of audio.
A new large open-weights multimodal model offers competitive transcription accuracy, increasing options for developers seeking alternatives to proprietary APIs.