Kimi K3 represents a breakthrough as the first open-weights model that achieves frontier-level performance specifically in agentic coding tasks, featuring a 3-trillion parameter scale with a Mixture of Experts (MOE) architecture (896 experts, 16 routed per token), native multimodal vision capabilities, and a 1M context window; while it excels in coding benchmarks like Terminal Bench 2.1 and web development, it lags behind in general chat and hallucination rates compared to models like Opus 4.8, making it ideal as a specialized workhorse for agentic coding use cases rather than a general-purpose chatbot.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Kimi K3 Explained!
Added:Okay, we need to talk about Kimi K6B because this is another deep seek moment. Probably bigger than that.
For the first time, we have an open weight model that is not quasi frontier.
But it actually is going to set the frontier itself.
When it comes to coding, it's probably one of the best model out there. On artificial intelligence index, this is the third ranked model surpassing Opus 48. Now, this model is really great when it comes to web development and UI design. You probably have seen some of those demos and that's not the focus of this video.
I am going to create a couple of videos testing this against other models, but I want to talk about something else.
And we're going to really focus on what exactly this means for the whole ecosystem because the Kimi models play a critical role in it. This is the biggest open weight model that we have seen.
This is a 3 trillion parameter model, which is enormous. But it also gives you a sense of what the size of some of these closed source models are going to be.
Now, with this release, Kimi is really breaking a lot of assumptions that industry experts had. People thought that open models are probably 6 to 12 months behind the frontier.
But with this, that gap has completely vanished. We are at the frontier right now when it comes to open models.
Which is kind of incredible.
However, there's a very important aspect that we need to discuss. We are at the frontier for specialized models, not general models. If you look at the performance even on the artificial intelligence index or something like the arena, it surpasses Fable class models on code or web dev arena.
So, you could arguably say that this is the best coding model at the moment.
But when it comes to general text generation like chatting, it is actually well behind the frontier.
Now, the thing is that Kimi or Moonshot has been really focusing on agentic coding or agentic use, and that is their specialty.
And this is where this model really shines. And this is a very important distinction to make because a lot of people just look at these aggregate results and think that it's overall the best model.
So, if you look at something like uh humanities last exam, which is not critical for agentic use cases, it lags behind Fable 5 by a large margin. In fact, when it comes to hallucination rate, it's actually worse than Opus 4 8.
However, when it comes to agentic coding, if you look at a benchmark like Terminal Bench 2.1, which is a lot harder than Terminal Bench 2, it surpasses Opus class models by a huge margin.
So, this might not be the best model to talk to. However, this is a workaholic model, and you can use it as a workhorse, which is what a lot of agentic coding use cases are going to need. Now, when the weights are available, this is going to be a favorite for a lot of companies to build on top of this.
We actually have seen the same pattern with previous Kimi models. It's uh ecosystem friendly, and it's favorite when it comes to further post training.
So, we're going to be seeing a lot of models which are built on top of the Kimi models because we have seen an industry pattern for this. If you look at uh models like Composer 2.5, those are post trained on top of Kimi K 2.5.
So, in general, the Kimi models are really strong when it comes to agentic coding. And the most interesting thing is that Cursor was able to push its performance even further.
And there is a lot of speculations that this model is probably mid-trained, which means there is a lot of juice that you can squeeze with further post-training.
Now, Kimi models has been special. So, even Cognition or previously Wren surf used Kimi K 2.7 as a base for their Sweet 1.5. In fact, I covered Inkling from Thinking Machines in my previous video. They used Kimi 2.5 for synthetic data generation to train their Inkling model. So, Kimi creates some really amazing models, which are often mid-trained with mid-RL, and they are really great candidate for further post-training. So, if you have compute to post-train an a 3 trillion parameter model, I think this is going to be a really interesting base for some really interesting models.
Okay. So, architecturally, this is much sparser compared to some of the previous models that we have seen. It has 896 experts, and a single token is routed through 16 out of those 800 and per token.
Now, we have seen a lot of people or companies are experimenting with MOEs because they are a lot more friendly for inference, but architecturally, it's a new architecture, and it has native vision capabilities, so it's multimodal in nature from the ground up, and it has a 1 million context window.
So, it's going to be more than sufficient for a lot of agentic coding use cases.
We don't know the parameter count of some of the closed-source models.
However, if you look at the performance of this model, you can make a pretty good estimate. I would say Opus is probably a similarly sized or maybe slightly bigger, but not substantially bigger than this.
Now, let's talk about the part that a lot of people are not going to like it because with frontier level model, they have a frontier level price increase as well.
You are essentially looking at Sonnet level pricing, which gives you performance close to Opus 4.8 or surpasses that. In some cases, you are getting Fable 5 level performance at the cost of a Sonnet model. Now, price per million token is just one part of the story. More importantly, how many token it takes to complete a task. And it seems to be pretty efficient as well, especially on the artificial intelligence index. It actually achieves similar score at a much lower cost compared to something like Fable 5 or even Opus 4.8. So, this is going to play in favor of Kimi or Moonshot because in the past, their model has been a lot more token intensive compared to some of the proprietary models. So, there's a very interesting study from Databricks, which showed that cheaper models in general are not cheaper when it comes to per task usage. They showed that Sonnet level models are usually a lot more expensive compared to something like Opus level model because Opus models are a lot token efficient compared to Sonnet models. And this token efficiency has been one of the primary focus area for OpenAI. They have been creating some really amazing token efficient models.
So, with all the conversation around our token maxing and cost of running these frontier models, it's really great to see that Kimi is actually focusing on token efficiency.
So, this may not be as efficient as some of the other offerings, but directionally, seems like they're doing a really good job. Okay, so what exactly does this mean for you as a end user? I don't think this changes much if you are interested in local models for you because you can't run a two or three trillion parameter model on consumer hardware.
This model is really geared toward enterprises who are going to be able to post-train this for their own purposes.
If you're going to be using it, you're probably paying less compared to some of the other closed-source uh models, but your data again is not private. You are already at the whims and desires of the inference provider. So, keep that in mind when you are using any model, whether it's proprietary closed-source or open-source. However, this is going to drive the ecosystem. There are going to be companies like the Cursors, Cognitions, and Thinking Machines of the world who are going to make really great use of it and essentially are going to provide better products for end users.
Now, the biggest advantage of this is going to be that this is going to force other labs to probably reduce their prices because you will have a lot more options in order to use frontier-level models.
Now, in general, looking at uh some of the outputs that this model is able to generate, this has some really great taste. And if you're interested in that testing videos, I am planning on creating a couple of videos comparing the generations directly to Fable 5, GPT 5.6, Soul, and also with the current best open model, uh GLM 5.2. So, if you are interested in those, make sure to subscribe to the channel. Anyways, I hope you found this video useful. Thanks for watching, and as always, see you in the next one.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

WOW! Judge TURNS THE TABLES on Trump in His OWN $10B LAWSUIT!!!
MeidasTouch
197K views•2026-07-23

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Steam and Xbox Just Dropped The Hammer On PlayStation
OhNoItsAlexx
9K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23