The release of Kimi K3, a 2.8 trillion parameter open-weight AI model by Moonshot AI, demonstrates how open-source AI models can compete with closed proprietary models like Claude and Gemini, with the model achieving top rankings in frontend coding benchmarks while offering free weights for download, thereby creating new competitive dynamics in the AI industry and challenging existing pricing models.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
China Just Dropped a 2.8T Parameter AI: Goodbye Claude?
Added:While Google was busy missing its third launch deadline in a row, a Chinese lab shipped the largest open model in history.
2.8 trillion parameters.
It beats Claude at front-end coding.
And in 9 days, the weights go free for everyone.
I build with Claude every single day.
So, I spent today finding out if my stack just became obsolete.
And hidden in this launch, there's one 25-cent detail that tells you everything. 9 days. That's how long until anyone can download this thing.
Wednesday night, Moonshot AI out of Beijing drops Kimi K3. The headline number, 2.8 trillion parameters. The largest open weight model ever built.
For scale, that's over four times the size of anything the open world had before. Now, it doesn't run all of that at once. It's a mixture of experts.
Picture a building with 896 specialists inside. For every word it generates, only 16 of those experts wake up and do the work. That's how something this big stays fast. It reads images natively, and it holds a 1 million token context window. That's roughly 10 novels in its head at once. Under the hood, there are two brand new tricks. Kimi Delta attention and attention residuals.
Translation. It remembers long documents better while burning less compute. But, none of that is the actual story.
The actual story is the license.
This thing isn't locked in anyone's cloud. July 27th. That's the date Moonshot publishes the full weights.
9 days from now, anyone on Earth can download a frontier model.
For free. And the timing?
The timing was an assassination. Because do you know what was supposed to happen the very next morning? Gemini 3.5 Pro.
Google's flagship. The launch the whole industry had circled on the calendar.
July 17th. It never came. Third missed deadline in a row. June slipped.
Early July slipped. and now reporting says it's delayed by months. The rebuilt model reportedly still hallucinates too much, and it's losing to GPT 5.6 on internal benchmarks.
Google's biggest launch of the year, and the lights just went out. So, picture that night, Google's war room going dark, and across the Pacific Moonshot hits publish. You could not write better television. One night, one company shipped the biggest open model ever.
The other missed its third deadline for the model it already announced twice.
That contrast is the whole AI race in a single frame. Okay.
But, is it actually good? Let's do numbers, and let's be honest about them, because most videos today won't be. The benchmark that matters most to me is GDP Eval.
It measures real work across 44 occupations in nine industries, not puzzles, jobs. Gemini K3 scores 1687.
Third place overall, ahead of Claude Opus 4.8, behind only the two absolute frontier models from Anthropic and OpenAI. Read that again. An open model weights about to be free just beat Opus 4.8, the model that was Anthropic's flagship 6 months ago. And here's the part I respect.
Moonshot themselves say it.
Overall, K3 still sits behind Claude Fable 5 and GPT 5.6.
They printed that in their own release, but there's one arena where it isn't behind anybody, front-end code.
The leaderboard where humans vote on which model builds better interfaces.
Number one, 1679 points, above Fable 5, above Seoul.
Within 24 hours of launch, a Chinese open model topped the coding chart everyone watches. Zoom out to the overall intelligence index, and it lands third, basically tied with Opus 4.8 and GPT 5.5. And it got leaner doing it. 21% fewer output tokens than the last Kimi.
Smarter and more efficient. That one stung. Front end is my bread and butter.
So, before we panic, let's look at what it's actually like to use this thing because the best review of it wasn't a benchmark at all. Simon Willison, one of the most trusted independent testers in AI, ran his famous stress test on it.
He asks every new model to draw a pelican riding a bicycle as SVG code.
Silly on purpose, revealing every time.
Kimi's pelican came out genuinely good.
But, look at what it took.
13,241 reasoning tokens of internal thinking to produce a 3,400 token answer. One drawing of one pelican, 25 cents.
Remember that number because we're coming back to it. He found something else, too.
Say just hi to Kimi K3 and the token counter reads 86.
Which means there's an invisible 85 token system prompt riding along on every single message you send. Nothing sinister. Every provider does something like it.
But, it's the perfect reminder. With any hosted AI, you're never quite alone in the room. His full write-up is the sanity check this launch week needed.
Not cheerleading, not panic, just receipts. All right.
Now, the pricing page.
You know this is my favorite part.
Everyone else reads the headline. We read the fine print. $3 per million tokens in, 15 out. That is Claude Sonnet territory and it makes Kimi K3 the most expensive model any Chinese lab has ever shipped. Compared to the big guys, it's still roughly half the per task cost of Opus 4.8.
Cheaper, yes. Cheap, no. And here's the trap in that math. Remember the pelican?
13,000 thinking tokens you pay for, but never see. A quote-unquote cheap model that thinks out loud this much can quietly outspend the expensive one. We learned that lesson with Sonnet's tokenizer. Same trap, new wrapper. The pattern here has a name and you already know it, deep seek.
18 months ago a Chinese lab shocked everyone with cheap frontier performance and wiped $600 off Nvidia in a day. That day, trading floors learned that AI modes can evaporate between market close and market open. Same script this week.
AI valuations wobble. Analysts started asking whether the compute mode is real and Washington noticed. David Sacks, who advises the White House on AI, called it a warning shot.
America risks losing the race through its own red tape, he says.
Blocked data centers, stacked regulations. There's a nastier question underneath, too.
Critics keep asking whether some Chinese models learned by imitating American ones.
The labs call it distillation.
The lawyers will call it something else.
Whatever you think of the politics, the technical fact stands. Export controls were supposed to make this impossible.
Moonshot built it anyway. Now the caveat this story needs.
Independent testers found Kimmy's previous model gave riskier answers on dangerous content prompts than its American rivals.
The K3 safety report card doesn't exist yet and your data goes through Moonshot servers under Chinese jurisdiction.
For a hobby project, maybe fine.
For client work, that's a real conversation. Filify runs on client documents all day, contracts, personal data.
I don't get to be casual about where that goes, which is exactly why the July 27th weights release is the real event.
Local weights mean your hardware, your data, no server in Beijing.
That's when this stops being a demo and starts being infrastructure. Nine days from now that calculation flips for anyone with a beefy enough machine or a rented GPU box. So here's my honest builder verdict after day one.
Front-end prototypes and UI work, genuinely tempting. It's number one there for a reason.
Long agentic jobs and anything with client data, I'm staying on Claude.
And watching that thinking token meter like a hawk either way. Three things to take away.
One, the open weight frontier is real now. The gap to the closed labs is months, not years. Two, pricing pressure just got a new engine.
When a free to download model beats last year's flagships, every API price in the market has to answer for itself. Three, watch reasoning tokens.
The new pricing games aren't on the price page. They're in the thinking meter. Total honesty, today I'm not switching. But for the first time since DeepSeek, I opened my stack and asked the question.
That's what a real launch does. Here's the twist though, while everyone's arguing about China tonight, the model Kimmy is chasing, Claude Fable 5, has a completely legal free tier that most people don't know how to use properly.
Monday, I'm showing you exactly how to run the best model in the world without paying.
Subscribe so you don't miss it.
I'm Dex. Go build something.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

Ben Crump dealt MAJOR BLOW after His Own Nolan Wells Autopsy FACT CHECKS him
DeVoryDarkins
50K views•2026-07-23

Gremlin Arrives… While Dorothy May Takes Another Step Forward
The-moons
10K views•2026-07-23

Trump War Chief SCREWS UP by Posting Video Leading Judge to ORDER an EXPLANATION!!!
LegalAFMTN
110K views•2026-07-23

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23