Inkling, a 975B parameter Mixture-of-Experts model by Thinking Machines (founded by Mira Murati, former OpenAI CTO), demonstrates that honest positioning—openly admitting it's not the strongest model—can be a smart strategy. The model uses 41B active parameters per token (256 total experts), supports native multimodal reasoning (text, image, audio), and offers controllable thinking effort with 7 settings. Despite matching Nemotron 3 Ultra at one-third token cost, it cannot run locally (requires 2TB VRAM) and lands mid-pack on benchmarks. The real value lies in its Apache 2.0 license, which enables hosting, fine-tuning, and commercial use without restrictions, contrasting with many 'open' models that bury limitations in custom licenses.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Mira Murati's New AI: Why Inkling Isn't Trying to Be
Added:A brand new lab shipped a 975 billion parameter model.
Then admitted it isn't the strongest.
That concession is the entire strategy.
This is the most honest artificial intelligence launch of the year.
The lab is Thinking Machines, founded by Mira Murati, the former Chief Technology Officer at OpenAI.
On July 15th, it shipped its first foundation model, Inkling, and put the full weights in the open.
975 billion parameters. Anyone can download.
Two questions hang over it, and neither answer is the benchmark.
Why open by admitting it isn't the best?
And if the weights are open, why can't you run them?
Underneath Inkling is a mixture of experts.
A giant model that only uses a slice of itself at a time.
Picture 256 specialists.
A word comes in, a router wakes just six plus two that are always on.
Everyone else stays dark.
On paper, 975 billion parameters.
Doing the work on any given token, only 41 billion.
A 975 billion parameter library with a 41 billion working brain.
That's how a model this size runs at the same speed. You never light up all 256 at once.
And it scans a full million tokens in one pass.
A whole code base or a long recording in a single shot.
Inkling doesn't stop at text.
It trained on 45 trillion tokens of text, images, and audio.
All three go straight into the same reasoning core.
Most models bolt vision or audio on afterward as a separate adapter.
Inkling takes image patches and raw audio natively.
Reasons over them together.
Answers in text.
Hand it a screenshot, a chart, or a voice clip.
It treats them as one problem, not three.
For a small team, that's one model to run and one bill to reason about.
No stitching a vision service and a speech service around a language model.
But reasoning has a price.
Every thinking model burns extra tokens to work a problem through and you pay for everyone.
Inkling puts that on a dial.
You pick the effort from none up to max.
Seven settings.
A quick look up stays cheap.
A hard proof gets the deep pass.
And the efficiency number is the real headline.
On terminal bench thinking machine says Inkling matches Nemotron 3 Ultra using about a third of the tokens.
For a small team watching its inference bill, fewer tokens at the same quality beats topping a leaderboard.
Here's the catch, the headline skip.
You cannot run this on your own machine.
Open weights does not mean download and run.
The file is the easy part.
To load Inkling at full precision, you need about 2 terabytes of VRAM.
Eight of Nvidia's B300 chips or 16 H200s.
Even the compressed NVFP4 version wants around 600 GB.
Your laptop has maybe 8 to 96 GB total.
So it runs on data center Blackwell and Hopper GPUs, the B300s and H200s.
Not your laptop and not a single gaming card.
Open weights got you the keys.
They did not get you the forklift.
So how do you actually use it?
You do, just not the way you assumed.
The inference providers you already reach for, Together AI, Fireworks, Model, Databricks and Base 10 are hosting Inkling from day one.
The open weights become a normal API call on the stack you already pay for.
Cheaper, faster, open model inference with no GPU purchase.
Want to shape it to your own domain?
Thinking Machines runs a managed fine-tuning service Tinker that trains on their hardware, not yours.
They launched it at half price with a free playground. And if 975 billion is too heavy, there's Inklings more.
A 276 billion parameter preview. 12 billion active tuned for coding.
Treat the big model as a blank base you carve down to your own problem.
Quick word about Hostinger.
Four ways to run AI agents on your own infrastructure.
AI agents, connector, open claw, and Hermes agent. Use code DIY smart code for an extra 10% off at hostinger.com.
Okay, back to the video.
Now the benchmarks.
Inkling isn't on top, and that's on purpose. Against eight other frontier models, it lands mid-pack almost everywhere.
On the AIME math test, 97.1%.
GLM 5.2 hits 99.2.
GPT 5.6 still 99.9.
Once we bench 77.6.
Behind Kimmy K 2.6 at 80.2.
Well behind Claude Fable 5 at 95.
So it loses to the open models, too.
Why cover it?
Here's the reframe.
Thinking Machines told you up front it isn't the strongest.
The score was never the pitch. The pitch is an Apache licensed base you can host, fine-tune, and reshape. A different product than a number on a chart.
Which is why the license beats the leaderboard here.
Plenty of models call themselves open, then bury a custom license that blocks commercial use or gates how you deploy.
Inkling ships under plain Apache 2.0, confirmed right on the model card.
Build on it, ship a product, no lawyer on the fine print.
And that matches the honesty from the very first line.
A lab that tells you where it loses is easier to trust on where it wins, like native audio and controllable cost. The honesty is the whole product.
Real quick, the dynamist.ai community is where I go deeper on this stuff.
So, here's the question.
Inkling openly admits it's not the strongest model.
Is that honest positioning a smart strategy or a cop-out?
Tell me below and subscribe for more open model news.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

we're almost finished the house (ep.125)
JennaPhipps
347K views•2026-07-22

We Finally Know Where Saturn’s Rings Came From
astrumspace
79K views•2026-07-22

BIG BET: Cathie Wood goes ALL IN on Elon Musk
FoxBusiness
89K views•2026-07-22

MIC DROP: Smithsonian Director Called Out For Woke Propaganda
TheAmalaEkpunobi
37K views•2026-07-23