Quick Facts
- Inkling is a 975-billion-parameter mixture-of-experts model with 41 billion active parameters per task, trained on 45 trillion tokens across text, image, audio, and video.
- The model is released under the Apache 2.0 license and is currently the largest open-weights model from a US-based AI lab.
- Thinking Machines monetizes through Tinker, its fine-tuning developer tool, not through Inkling itself.
Thinking Machines Lab, the AI startup founded by former OpenAI CTO Mira Murati, released its first public AI model Wednesday. The model, called Inkling, is available as open weights on Hugging Face under the Apache 2.0 license, allowing developers to download, inspect, and fine-tune the full codebase without licensing fees.
Murati founded Thinking Machines in late 2024 after departing OpenAI, alongside industry veterans John Schulman and Barret Zoph. The company built Inkling in roughly nine months, the lab says, marking its first public proof point after more than a year of work largely outside the spotlight.
What Inkling Is
Inkling uses a mixture-of-experts architecture with 975 billion total parameters. It draws on about 41 billion of those for any given task, a design that keeps large models faster and cheaper to run. The model was trained on 45 trillion tokens of text, image, audio, and video data, though current outputs are limited to text, code, styled artifacts, and structured data. Its context window reaches 1 million tokens.
Inference is live across Together AI, Fireworks, Modal, Databricks, and Baseten. Thinking Machines also released a preview of Inkling-Small, a lighter model with 12 billion active parameters built on the same training approach.
Benchmark Numbers
On SWE-bench Verified, a software engineering benchmark, Inkling scored 77.6%, beating Nvidia’s Nemotron 3 at 71.9%. On AIME 2026 math problems, Inkling scored 97.1%, edging DeepSeek V4 Pro at 96.7%. On VoiceBench, it scored 91.4%. On IFBench, a chat evaluation, it posted 79.8%, ahead of Claude Fable 5 and GPT 5.6 Sol in the company’s comparison table.
The company also says Inkling uses one-third the tokens of Nvidia Nemotron 3 Ultra to reach equivalent coding performance. All benchmark results come from Thinking Machines’ own reporting. Independent third-party evaluations have not been published as of this writing.
China’s leading labs remain ahead on pure reasoning and coding tasks. GLM 5.2 outperforms Inkling on those dimensions, according to the benchmark data Thinking Machines shared.
The Business Case
Thinking Machines is not monetizing Inkling directly. Revenue comes from Tinker, its fine-tuning developer tool. The company sells Tinker to customers such as hedge fund Bridgewater Associates, which used it to train a model on Bridgewater’s financial knowledge. That model scored 84.7% on financial reasoning tests, beating top proprietary models, according to a joint evaluation by both companies.
Murati said on X: “Our first model, Inkling. Trained from scratch, weights are open, fine-tunable on Tinker today.”
The strategy targets enterprises skeptical of closed-model costs and data exposure. Microsoft CEO Satya Nadella has warned that firms using closed models pay twice: once in fees and once by handing over knowledge embedded in their prompts. Palantir CEO Alex Karp made similar comments on CNBC, saying frontier tools from closed providers are too expensive and lack clear IP protections.
Thinking Machines is betting that enterprises care more about owning and shaping a model than accessing the highest-performing general-purpose one. Inkling combined with Tinker is the company’s pitch to run that enterprise learning loop. Nvidia CEO Jensen Huang backed the direction, saying Thinking Machines has “brought together a world-class team to advance the frontier of AI.”
Read more: Mira Murati’s Thinking Machines drops Inkling, an open-weights model anyone can access
This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.
