Thinking Machines Lab, the company founded by former OpenAI chief technology officer Mira Murati, released its first in-house model on 15 July 2026. Called Inkling, it is an open-weight Mixture-of-Experts model with 975 billion total parameters, 41 billion of them active for any given token, a context window of up to 1 million tokens, and native multimodal understanding of text, image and audio inputs while generating text output. The full weights are published on Hugging Face under an Apache 2.0 licence, and the company is explicit that Inkling is not built to top the leaderboards.
What is actually in the model
Inkling was pretrained on 45 trillion tokens spanning text, images, audio and video, and reasons natively across text, image and audio while producing text output. Its Mixture-of-Experts architecture largely follows that of DeepSeek-V3, with 256 routed experts and two shared experts, six of the routed experts active for any given token. Developers can set a thinking-effort level to trade accuracy against token cost. Alongside the main model, Thinking Machines previewed Inkling-Small, a 276-billion-parameter version with 12 billion active parameters, whose weights will follow once testing finishes.
The company's own positioning is unusually direct. It states that Inkling is not the strongest overall model available today, open or closed. Instead it is offered as a broad, multimodal base that organisations fine-tune on their own data through Tinker, the lab's customisation platform.
Where it lands on the benchmarks
On the independent Artificial Analysis Intelligence Index, Inkling debuts at 41, which makes it, at the time of writing, the highest-ranked open-weight model released by a US lab, ahead of Nvidia's Nemotron 3 Ultra at 38, Google's Gemma 4 31B at 29 and OpenAI's gpt-oss-120b at 24. That number also bounds the claim: 41 sits well below the closed frontier and below China's Kimi K3, which scores 57 on the same index. Thinking Machines' own tables show Inkling competitive with other open-weight models across reasoning, coding and multimodal tasks, but those figures were produced partly on the lab's own harness with self-reported numbers for rival models, so they are company-reported rather than a like-for-like comparison.
The wager: customisation over raw capability
Murati's argument, echoed across the launch coverage, is that many enterprises care less about renting the single smartest general model than about running one they can adapt to their own workflows and data. Inkling is available for fine-tuning on Tinker at a temporary 50 per cent discount, and for inference through TogetherAI, Fireworks, Modal, Databricks and Baseten. The lab also disclosed that Inkling's post-training was bootstrapped in part on synthetic data generated by open models including China's Kimi K2.5, a sign of how freely capability now moves between labs. According to Thinking Machines' own safety evaluations, Inkling achieved the strongest built-in safeguards among the open-weight models in its comparison on the adversarial FORTRESS benchmark, at 78 per cent, and scored 98.6 per cent on the StrongREJECT refusal test.
Key Takeaways
Inkling, released 15 July 2026, is Thinking Machines' first model: an open-weight MoE with 975 billion total and 41 billion active parameters, up to 1 million tokens of context, and text, image and audio input, under an Apache 2.0 licence.
The company states Inkling is not the strongest model available and pitches it as a customisable base for fine-tuning on its Tinker platform.
It debuts at 41 on the independent Artificial Analysis Intelligence Index, the top US open-weight result, but below the closed frontier and China's Kimi K3 at 57.
A preview of Inkling-Small (276 billion total, 12 billion active) was shared; its weights follow after testing.
Company-reported benchmark tables were run partly on the lab's own harness; the independent anchor is the Artificial Analysis index score.