Kimi K3, a 2.88 trillion parameter model from Moonshot AI, represents a significant breakthrough in open-weight large language models by achieving performance comparable to or exceeding frontier models like Claude Fable 5 and GPT-5.6 in multiple benchmarks (including coding, software engineering, and agentic tasks) while maintaining competitive pricing at $0.30/MTok for cache-hit input. The model's native multimodal architecture and Kimmy delta attention mechanism enable 6.3x faster decoding and 25% higher training efficiency, demonstrating that open-weight models can compete with closed-source alternatives from major AI labs.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Kimi K3 BEATS Fable 5 and GPT-5.6! This changes the landscape
Added:Yesterday I made a video about how Kimmy K3 was kind of leaked and we saw early outputs from that model and surprisingly the model well not surprisingly I knew this was coming was launched today. We knew that this model was around the corner and when it would launch we were expecting big things from Kim K3. We're expecting according to early claims this model to be quite competitive. This model to actually match Fable performance as well sometimes even outperform it. And today it looks like that is definitely the case. So Moonshot AI actually published this. They published the actual blog and the model is live now. So you guys can actually access Kimmy K3 on kimmy.com. And what's important to note, the open weights will be coming on July 27, 2026. So once again, big win for a lot of us, especially people who are running the models on their own. So what exactly is new with Kimmy K3? So first thing, it's actually a 2.8 8 trillion parameter model with 1 million context and is natively multimodal. So multimodal features are not bolted on after this is a native multimodal model and we can see that has helped the model a lot especially in game development which I'm going to show you guys in a bit. There's also something that they've added which is Kimmy delta attention which enables up to 6.3 times faster decoding in million token context. In short, this basically means it's an efficient model.
Attention residue deliver 25% higher training efficiency at less than 2% additional cost. Another importance on efficiency. And this model is actually built for long horizon agentic coding and self- evvolving workflows. The self-evolving part they actually show a I guess an example or a demo where this model actually beats Fable 5. Obviously, these are Moonshot's own independent claims, but I'll show you guys those claims and you guys can be the judge of it when you actually use the model or over time as we see people actually working with this model and see if it holds up. Now, on the benchmark sides, I want to be clear, this model does not beat GBT 5.6 completely or you know, Fable 5 completely. There are some areas where it does, some area where it doesn't. But what's really important to know here is that this is an openweight model. This is an openweight model that's matching performance of the best models from Enthropic and OpenAI. So that is what I want you guys to focus on when you are looking at the benchmarks.
Not if it's number one, but how capable this model is. But on the get- go, this model basically beats Opus 4.8 across the board. So if you're an Opus user and if you are enthropic and you are you know waiting to release Opus 5 which we know is probably coming soon this might be the time because Kimmy has basically overtaken Opus 4.8. Now if you look at the coding benchmarks on the deep software engineering one Kimmy comes in at 67.5 which is above GPT 5.5 and Opus 4.8 and obviously a little bit below Fable 5 and GPT 5.6. So now what's really crazy to me is that on the Frontier Software and Engineering benchmark, the model comes in at 81.2, which is actually above GPT 5.6 soul. Same thing on the Kimmy Code Bench 2.0, which is an internal test. So maybe take that with a little bit of grain of salt. The model comes in at 72.9 again above Opus 4.8 and GPT 5.5.
And obviously surprisingly, GPT 5.6 6 souls performed worse than GBT 5.5 on this benchmark, which I think is interesting. But anyways, on Terminal Bench 2.1, Kimmy K3 once again comes in at second, 88.3. And guess what the top score on that one is? 88.8. So about.5 of a difference from GBC 5.6. So, and then this I think is really crazy. On the program bench, Kimmy K3 actually comes in at number one at 77.8 8 and GBD 5.6 soul 77.6 which tells you that okay in some areas this model performs better than the frontier labs. Same thing with software engineering marathon. Kimmy K3 comes in on 42.8. This here's basically the long horizon coding angle that they are kind of getting at in the blog post. Now these are other tests. I'm not going to go through all of them but you can see even on general agents all maxed out on thinking efforts. Remember that this is all maxed out. the model achieves Kimmy K3 like third place on GDP val and then and the other ones automation bench which I think is important to highlight Kimk 3 comes in at first above GPT 5.6 soul and fable 5 same thing with browser comp the model comes in at first 91.2 too. And then on spreadsheet bench 2, Kim K3 comes once again number one. So knowledge workers good win for you guys as well because this model is good at spreadsheets. So good for them. And then for visual agents where it comes at, you know, understanding charts with tools and everything like that, the model comes in second. And then same thing with the Zero Bench with tools, the model comes in at second. So yes, this model is competitive and it was able to, you know, achieve a lot of this. But I think some of the reason why the model was able to do all of this is because of the architecture behind this new model.
And this is where it gets a little bit complicated. But I think the main thing to understand over here is that all of these improvements have made the model more efficient. And we can see that because we see an approximate 2.5 times improvement in overall scaling efficiency compared to K2 which is their older model. And what are they actually calling this new architecture? So K3 is actually built on Kimmy delta attention KDA and attention residual attention rest which are two architectural updates designed to specifically improve how the information flows across sequence length and model debt. Now that sounds complicated. It probably is but it shows you that the lab is actually putting in research behind all of this. And one thing that you guys are probably aware of is that if you have been tracking research papers in the AI space, a lot of the times they come from labs that are in China or they come from people who are tied to the Chinese labs, especially Deepseek. And we see that again with Moonshot right now, which tells you that, you know, sometimes you don't need to just throw money at a problem. You need to throw some intelligence, which is what China has been doing. And we can see that clearly because of the efficiency improvements in the model. 2.5 times improvement over the last model which is huge. Now these are some internal knowledge work benchmarks and obviously these are Kimmy's internal benchmarks. So everything you see here I would you know maybe like be a little bit skeptical but obviously Kimmy K3 comes out as number one on the online Xbench deck bench basically creating slide decks finance bench. So they are also trying to target knowledge workers. They're just not focused on developers. They want knowledge workers to be using their models as well. And we can see that especially with the last two benchmarks which is creating slide decks as well.
Finance bros this model might be good for you as well. Now this was a self-evolving task that Kimmy wanted to highlight but basically the summary is that basically the model had to work on a problem and for over 15 hours of non-stop iteration K3 designed a novel two-phase kernel algorithm fuse kernels while preserving numeric and reduce forward backward time from 28.36 milliseconds to 114.4 4 milliseconds and they're saying that K3 and Fable 5 with potential fallbacks meaning the model might have you know called back to a lower model Opus 4.8 Okay, reach similar performance, but according to them, K3 improved faster per iteration. As I said, these are all internal benchmarks, but it still shows you that this model is quite competitive and capable even on self- evolving, which is a capability of more and more new models. We're seeing that K3 comes on on top, which tells you that a lot can happen in the AI space very fast, and we can see that clearly with how much Kimmy K3 has improved. Now one of the areas that I mentioned was that multimodal is built into the model and that has helped the model produce a lot of strong 3D reasoning coding and vision capabilities and because of that we are seeing like game development is actually an area that this model excels at. They actually even show video editing in their demos in the blog post which I'll link to if you want to look at the full blog post. But it tells you that this model is not only good at coding but this model is getting better at more tasks than just you know being really good at solving something in the terminal for example like maxing it out in terminal bench because a lot of time in the past people used to criticize the models from China for benchmaxing meaning that they would train the model to perform really well on a specific benchmark but we can see that Kimmy K3 excels in a lot of areas we see that with vision capabilities we see that with knowledge work we see that also with programming And now we're also seeing it with game development.
Obviously, they're all related a little bit, but these improvements tell you that Kimmy K3 is probably one of the biggest release in my opinion right now in 2026. Not because this is the best model. Obviously, in some areas, Fable 5 and GPC 5.6 so still are better than this model. But based off what we know about Moonshot, based off what we know about openweight models, this model basically told all the Frontier Labs that hey, you guys can basically produce the best models. Your models are expensive. We can do the same at a fraction of the cost and here's the proof. And they just proved that with KBK3.
Now on Ellarina which is basically you can think of it people actually working with the model and voting for the model on the frontend code when it comes to design capabilities basically the model actually comes in at number one now Marina the reason why this is important is because this is people actually voting for this so it's not going to be rigged by like obviously moonshot AI this is just people using the model and people voting for it and we are seeing that on Elmarina this model is better than cloud Fable 5 GP 5.6 sooul and GLM 5.2 from obviously China as well at front-end coding. So that tells you this model is quite competitive. Obviously we're going to see more benchmarks like these come up over the next couple of days and we might see people actually working with this models and things could change over time if the model kind of gets nerfed. Hopefully doesn't. I don't think Moonshot has a history of doing that. But this tells you once again that not only is the blog post talking about it, people are realizing the gains as well. Now, a piece I want to basically touch at is what does this kind of cost us. This is actually the same price and is pretty much on par with Sonnet 5, which is insane to see because you're getting a model that in some areas is better than Fable 5 and GPT 5.6 six and it's the fraction of their cost which tells you like okay there has been some serious change up in the whole AI space we kind of operated with a duopoly before I would say at least last couple of weeks maybe a couple of months as well where it was just open AI and enthropic now Kimmy K3 came out of nowhere I want to say moonshot actually cuz that's the lab behind it came out of nowhere and basically overtook them and also just changed the whole landscape definitely put pressure in my opinion back on OpenAI Anthropic because now they have to answer why their models cost so much when a model like Kimmy K3 exists. One thing as users is that it's clear that competition is really important in this space. Kim K3 is going to force Enthropic and OpenAI to re-evaluate how they're producing the models. They might have to start focusing more and more on efficiency. We kind of saw that with GPT 5.6. GPT 5.6 6 soul matched Fable 5 performance and it was more efficient than that model and that's what we saw with OpenAI kind of taking that angle with GPT 5.6 soul but the upcoming models which we know that Enthropic is working on Opus 5 as well open AI is working on GPT6 those models will now have to produce obviously better results than Fable 5 and GPT 5.6 six soul and basically be either the same price as they are or cheaper. Then only people will realize that they are heading in the right direction because Moonshot AI is doing that. They're producing capable models and the cost is quite reasonable.
But that's it for today's video. Make sure you guys are subscribed to the channel, follow our new newsletter as well at universeai.behive.com beehive.com as well as subscribe to the main channel World of AI and support us on X by following the universe of AIZ as well. Until then, I'll see you guys in the next video.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

WOW! Judge TURNS THE TABLES on Trump in His OWN $10B LAWSUIT!!!
MeidasTouch
197K views•2026-07-23

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Steam and Xbox Just Dropped The Hammer On PlayStation
OhNoItsAlexx
9K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23