Fireship brilliantly deconstructs the xAI hype, revealing that MoE is essentially a clever engineering hack to manage the computational gluttony of massive models. It’s a sharp reminder that true innovation often lies in efficiency rather than just raw scale.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
This $12 billion startup finally shipped something...
Added:Two years ago, Mira Murati did something we've all fantasized about. She woke up, decided she'd had enough of her boss, and quit her job as CTO of OpenAI with no backup plan and nothing but a vague desire to create time and space for her own exploration. If you or I did that, we'd probably end up coding HTML for food down on Market Street. But when someone like Mira does it, A16Z shows up with a $2 billion investment into any company she decides to build. The company was eventually Thinking Machines, and until last week, it's been known for being valued at $12 billion before having actually shipped anything.
But that may have now just changed because they just released Inkling, a fully open weights model trained from scratch that can see, hear, reason, fine-tune, and potentially touch things in the real world. In today's video, we'll find out how it was built, what makes it interesting, and learn why Thinking Machines openly admits that it's not the best model in the world, and why that's the point. It is July 20th, 2026, and you're watching The Code Report. Last year, while the rest of us were spending time getting to know our newest Soham Parikh, Mira was quietly pulling off the greatest talent heist in Silicon Valley history. When she left OpenAI, she walked out with co-founder John Shulman, VP of research Barrett Zoph, and enough senior researchers to fill a Silicon Valley polycule. A year later, the dream team shipped Tinker, an API that lets you fine-tune open weight models without managing your own infrastructure or losing control over the training process. When it was released, ML researchers called it cool, which I imagine was a bit underwhelming if you'd just given them enough money to buy a Costco hot dog for one out of every six people on Earth. But that may have changed last week with the release of Inkling, a mixture of experts model with 970 billion total parameters. The trick is that it doesn't use all of them at once. Every token gets routed to a small subset of specialized experts, so only 41 billion parameters actually fire per token, giving you the most intelligence of a massive model at a fraction of the compute. It was pre-trained on 45 trillion tokens of text, images, and audio, handles a 1 million token context window, and unlike other open models that ship with a license written by feral lawyers from Meta, the weights are Apache licensed and sitting on Hugging Face right now.
On the Trust Me Bro benchmarks, it gets mugged by Fable 5 and GPT 5.6 Soul, and lands somewhere mid-table among the open Chinese models. Now, that's not great, but the timing was even worse. One day after Inkling dropped, Moonshot announced Kimik 3, a 2.8 trillion parameter monster that, if you squint, goes bar for bar with Fable 5, and got so popular that it even burned through their compute. And so, in a sense, the team spent 2 years cooking up the most average model in existence, but that was kind of the point. Instead of competing on raw intelligence, Inkling has a dial called thinking effort that controls how hard the model cooks. You crank it down and you get cheap, instant answers. You crank it up and you get Ashley Tisdale's hit single from 16 years ago, with the model matching Nemotron 3 Ultra on Terminal Bench while using a third of the tokens. And when you're running an agent millions of times per day, that adds up. And the flagship demo was pretty cool. They connected Inkling to Tinker and asked it to remove its own ability to use the letter E. From there, the model wrote its own training script, generated its own data, ran the job, then loaded its own new weights that essentially lobotomized itself. They also trained it on something they call epistemics, which is a fancy Greek word for knowing when you're full of crap.
Basically, Inkling was rewarded for admitting what it doesn't know instead of guessing with confidence. As a result, it's now one of the best models in the world at forecasting future events, beating GPT 5.5 and Opus 4.8, and when it's unsure, it'll admit it.
Inkling is also the Helen Keller of models, not in the sense that it's bad at laser tag, but that it can't actually see or hear, at least not in the same way other models do. Normally, images and audio get run through separate encoder models that translate them into tokens first. Inkling skips all of that and processes raw audio and pixels directly. All that reinforcement learning had a side effect though. Is somewhere past 30 million training rounds, the models inner monologue started dropping words to save tokens.
It going from proper English to caveman speak like we need determine eigenvalue problem. Why waste time saying lot word when a few word do trick. So no, Inkling is not going to dethrone the Frontier Labs and Thinking Machines knows that.
The real play is to hand you a mid model for free and then charge you to fine tune it on Tinker into a specialist that destroys your one specific problem. But if you need a specialist today to handle things like auth, billing and user management, you need to check out Clerk, the sponsor of today's video. Your agents can build an entire app over the weekend, but if you somehow manage to get real users who want to pay you real money, you're the one stuck reading the Stripe webhook docs at 2:00 a.m. That's why Clerk built a CLI, agent skills, and an MCP server that let your coding agent drive their whole platform. You can set up production-ready auth, billing, and user management with just prompts. I told my agent to handle sign up and sign in and 13 seconds later I had auth pages, middleware, and a passing build.
From there, it added basic and premium paid plans and I never wrote a line of payment logic myself. When I asked for a list of my users, it pulled all eight of them straight from the API. A Clerk lets you go from dev to production without ever visiting their dashboard and you can try it out for free with the link below. This has been the Code Report.
Thanks for watching and I will see you in the next one.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

WOW! Judge TURNS THE TABLES on Trump in His OWN $10B LAWSUIT!!!
MeidasTouch
197K views•2026-07-23

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Steam and Xbox Just Dropped The Hammer On PlayStation
OhNoItsAlexx
9K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23