Kimi K3’s benchmark success shatters the narrative that Chinese AI is merely derivative, proving that open-source models can now achieve frontier-level performance at a fraction of the cost. This shift signals a new era of global competition where technical parity is real, but geopolitical trust remains the ultimate bottleneck.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
This Model is Better than Claude and ChatGPT
Added:New day, new model.
>> New day, new model.
>> You've been busy, man. I mean, gosh, this is like the fastest uh pace of model releases I think we've seen in a while.
>> It's incredible. Six models in the last uh seven days. It's been Gro 4.5, Kimmy K3. You know, previously we had GLM 5.2.
It's the the the new series of models, the GPT 5.6s, the Soul, you know, the Luna, the Terra. Uh it's been an extraordinary piece of model development and it's left us with a Chinese open source model on the top of the Koda >> and it's I mean it did it come out of nowhere here were we expecting we'll talk about the reviews in a second. Were we expecting this to be this good initially upon reviews?
>> Well you know there were there have been murmurss of the great performance of Kimmy K3 for quite a while. I'm not and but I'm not sure that people expected it necessarily from Moonshot. I will tell you what has been the trend which is that over the past year you've seen open source and closed source models had a big gap and then it started closing and closing and closing and then they were kind of tracking with clos with open source models just a few months or weeks behind uh closed source models and people were saying they're never going to crack it and it's because of distillation. They're just distilling, distilling, distilling. And so fundamentally, when you distill a model, it degrades. And so you're never going to get better performance. But this is the first time we're really seeing the narrative that breaks that mental model that says that the Chinese labs might actually just be really good at developing models and not just distilling American intelligence.
>> Right. So you run one of the most popular uh platform of benchmarks uh that is widely cited in AI. What is your data telling you about how good Kimmy K3 is?
>> Well, the way our platform works is that we have a a user base of tens of millions of people that are coming on Arena to use AI for their real workflows. They're agentic workflows.
They could do singlethreaded conversations that are, you know, tens or hundreds of turns long. They're getting to do, you know, workflow automation. They're doing coding.
They're doing math. They're doing all these sort of economically valuable tasks. And we take all those tasks and then turn them into a benchmark. And one of the benchmarks that we released for Kimmy K3 has been code arena. And specifically, Kimmy K3 is at the top.
It's number one in front-end code arena.
So, it's the best at the sort of lovable vibe coding use case. Now the best model in the world for that of course for to which hundreds of millions of dollars in inference spend will acrue is >> ahead of ahead of 5.6 ahead of fable too >> ahead of fable which is absolutely remarkable.
>> So let's just pause here for a second just this is this is this is kind of you know fable was the model that really uh uh caused a bit of a reckoning in terms of cost. This is an open- source model which I mean that means it's it's significantly cheaper right am I am I correct or >> that's correct and I will tell you the gap between Kim K3 and Fable is not small it's a win rate of in the order of tens of percent. You know, let's say 10% is the difference in win rate between Kim Kate 3 and Fable 5, which is, you know, it's it's not the largest gap that we've seen on Arena, but it is substantial and indicates a real delta in the change in performance. And now that comes at the cost of a model that is a sonnet level cost. That's a double-edged sw.
What do you mean by that?
>> Oh, like claw sonnet. It's it cost the same amount of cloud sonnet. It's $3 input, $15 output token. same price as um Claude sonnet. Um so the the significance of this is actually quite interesting. First of all, of course, the raw performance is amazing, but also the fact that it's no longer the case that open- source model is, you know, 50% of the quality at 10% of the cost.
It's actually that it's frontier level performance at a reasonable cost that you would see for a mid-tier frontier model. And so I think the ecosystem is evolving perhaps in an uh unexpected way in the sense that open source is actually also quite expensive which speaks to the fact that um the margins are actually not that high on inference.
>> Right. So that was coding. I mean just give us a brief overview here. Other benchmarks uh what other benchmarks does Kimmy K3 really excel at based on what you're seeing?
>> Well you know it does sort of like uh it's definitely a frontier model on a lot of the static benchmarks. So if you're looking at things like Swebench and Tao and so on and so forth, it's sort of in the pack along with um Soul and with Fable, although not necessarily leading leading on some, not leading on others, but still right there in the competition uh in text arena, which is our benchmark for sort of chat GPT like capabilities with user interaction and so on. Kimmy K3 is top 10, but not top one. And in fact, it's not even the top model outside of OpenAI Anthropic. It's about tied with soul, but it is beaten by new spark 1.1 and the basically entire bench of claude models uh 4.6, 4.7, 4.6 thinking, 4.7 thinking and fable all better um chat models for chat interfaces. So right right now is for text is who's at the top right now that >> right now it's Fable. It's Fable 5.
Yeah.
>> So so Kimmy K3 from what I'm hearing from you.
I mean really uh I mean it's it's coding that has really impressed me of all the use cases you could use this for. That is the one that's standing out.
>> Um let me ask you this. Are there caveats to this at all? I mean being a top a benchmark is certainly one way of of measuring performance. Are there other considerations though around the way the model is is made uh you know how easy it is to use uh that you know sort of might tell us something about how it's actually adopted uh by businesses.
You know we I we have yet to see how it's going to be adopted by businesses.
too early to say, but we're going to have more to say on that over the next days and weeks about how steerable the model is and how companies are finding it easier and easier to use. Of course, the one thing that I will say is Kim K3 is quite a large model >> and so it's not like you're going to be able to run this locally on your laptop.
You are going to need a cluster. Um it's going to need quite a bit of infrastructure. um and I would expect that many businesses will be relying on thirdparty inference providers to give them the compute they need even to do the inference on this model.
Let me ask you about the uh comparison.
I mean look deepseat came out this was a model that was not developed in the United States obviously it caused big market swings. Uh people sort of asked the question hey uh do I need all the compute uh that I initially thought? I mean, do you see any similarities with this model and and the impact that Deep Seek had on the ecosystem? Are they two separate stories? What do you think?
>> Well, I think it will cause a reckoning in the capital markets. And the reason for that is that it brings into question what the uh dominance will be of the closed source models in a world where the there's a narrative violation against the uh the distillation story. Then what will happen is that people will say well open source models coming from China they're you know they have so many benefits businesses can incorporate them into their own infrastructure without relying on a third party service without worrying about privacy and data leakage without worrying about you know evil third party companies training on their data and stealing their businesses. There's so many reasons why people would want an openweight model and in a world where that can be done for free. Then the question comes why would we be paying for closed source models that are worse for our businesses that are less private and that allow these companies that want to be in every single business in you know FDE for every single vertical. Why would we be allowing them to witness our business and and see our data, learn from it so that they can improve and and uh and eventually one day steal our business? Why would we do that? And so then what's likely to happen is an acrruel of value to these openweight models from which of course the revenue model is is is quite different um and and worse than the closed source models which could then cause a collapse you know in the way that people are seeing the compute markets because of course the compute markets are driven by the optimism and the revenue of the closed source models so on and so forth. So it could cause a cascading effect. Does it does it matter that it wasn't developed in the United States?
>> Absolutely. It absolutely matters. And in fact, we should all, you know, from if we're patriots uh and Americans be rooting for some American open- source models to come out as well. Of course, it's fantastic that we see the Chinese developing these models and it's socially positive net net, but at the end of the day, because of the geopolitical risk, we are not going to see businesses comfortable building their entire stack on Chinese models. Do you think that the the the I I hear you and that is certainly something we hear from from companies as well. Do you think that that is becoming less of an issue though? Because if I think of the last years, if I if I think about >> No, just >> Absolutely not.
>> Okay.
>> Why do you think it's less of an issue?
>> No. No. Well, I I I'm ask based on um this is this is purely based on, you know, uh narrative observation. What I I I feel like people are more willing to to use models not developed in the United States. That that's just based on conversations.
>> I think the government the governments of both sides would disagree. What's what I see happening is actually more hawk hawkishness on both sides that the Chinese government is actually talking about export controlling Chinese models >> and the US government is also talking about banning Chinese models because neither side wants to incur the geopolitical risk and of course there's so much competition happening between the US and China. It is it is absolutely becoming higher and higher risk. We even have the US government banning US models from release, let alone Chinese models.
I think it is actually a huge topic, >> right? Let me ask you about another model, Anastasia. So, uh, thinking machines lab came out with their model, Inkling. Uh, what's been the reviews that you've been seeing on on that particular model?
>> Listen, Inkling is the first open-source model to be released. uh by Thinky, actually the first model period to be released by Thinky. And of course, it's incredible to see that we have a company that is leading the American open-source charge. The other contenders would be Neotron, RC, Reflection, Mistrol, of course, uh in Europe, and there's a couple other players. Um but England came out at the top of the US open-source race today. That said, it is not yet a frontier model by any stretch.
It is, if you look at the wider field of top 10, including Chinese models, it is number 10. Number 10. So, there's nine Chinese models above this. And so, they have quite a bit of room to climb. Now, if you know anything about thinking machines, and it's sort of, you know, common knowledge, it's not something that I know specifically because of our business, but they've had quite a bit of restructuring, so on and so forth.
people leaving, people coming in, restructuring their team. And so maybe to some extent it's impressive that they were able to do this in a relatively short amount of time since the current structure that they've had 6 months or so. That said, um there is a long long way to go and if you look at the field of closed source models as well as open source models at a number 37 model. So we still have, you know, there's a ladder to climb, but let's all root for them because that is we need a model within our shores to support the enterprises that want to avoid the geopolitical risk that we just discussed. If we don't have an American open source model, um businesses will simply uh be incurring greater risk and be less comfortable adopting AI. Is there anything about the AI ecosystem in China that has uh you know made them so effective at developing open-source models specifically? It feels like in the open source race I mean it's always been a question of of uh which models are coming out of China. Why do you think they have always succeeded in that particular arena? No pun intended.
>> Yeah. Well, I think that um the strategy is actually quite a good one, which is to diffuse the model everywhere and then become the best place for running that model and therefore uh gain resources and you know they need to have some kind of an asymmetric advantage the Chinese model providers over the US model providers um there and but they because they have a lot of drawbacks too. So they are absolutely limited by the amount of GPUs that they have. they are absolutely limited by the fact that there's there's uh legal regulations inside the US that keep you from shipping data to China. Um and so because of these facts, they need to have some way of attacking the market and it may as well be open source. I think that's also compounded by the fact that by there's a lot of central planning in China. Um and IP is treated differently in China than it is in the US and um perhaps it's there's there's less of a focus of IP in the in from what I know of the Chinese market than the US market. So those are all reasons why the open source strategy would be better for them. And then how do you monetize? Well, there's two ways of monetizing an open source model. The first is become the best place to run that open- source model. And then you can earn money through inference. That's what a lot of these companies are doing.
And the second is through licensing. So you say okay I'm not going to run my open source model but I'm going to partner with inference providers in the US that are going to run this open source model and if they reach a revenue of above X dollars in running my model then I'm going to do a rev share with them and that's going to be in the licensing of the weights and of course because we respect licensing laws here that is exactly what those companies are doing and that's going to acrue back to revenue for those companies. I want to ask you uh where we started the conversation. We talked about the pace of the release of models lately. It it feels like they've been coming fast and furious and we've been talking on the show about where AI research is at the idea that um continual learning is is very much uh something that people are looking forward to. Why do you think models have been getting released so quickly and and let's put competition aside? I mean there is a lot of uh incentive for these companies to put things out to continue to stay competitive but is there anything about the current state of you know what these models can do whether or not they can help develop other models that is sort of uh influencing the the pace of the release of of these new products >> of course I think you're getting at the concept of recursive self the idea that the models might be able to train themselves and so on and so forth I don't think we're quite there yet although We are even seeing in our business that the AI um that the existence of AI is actually making our employees more and more productive by the month. We're seeing productivity per employee actually rise even at Arena at our small scale. We're 70 people and we are still seeing that effect happening.
So I'm sure it's happening at the big labs as well. That said, I will say there's a funny dynamic, you know, first of all, you know, even competition aside with these labs, which is that the release cycles tend to be kind of bulky in the sense that you'll get a quiet week with a couple models here and there, and then you'll get a week with 10 models. And why is that? I think these companies are all hearing about each other. Oh, Kim K3 is going to release the following day. Oh, we know that we're not quite as good as Kimmy Crate 3, so we should probably frontr run them so that we get a little bit of good press. Imagine if Inkling and Kim Crate 3 switch, right? If they switch on the release date and Kim Crate 3 comes first.
>> In Inkling doesn't they don't release.
They're not they're not coming out with it.
>> Yeah. Exactly. They don't they don't look quite as good and so it creates this dynamic of a mad rush. Everyone hears that everyone else is releasing and so everyone is competing to be first, especially the players that know that they're on a weak foot. And not to say that that Thinking Machines is. I think that the model's great. It's the number one US open source model and again we should all be rooting for that.
But it's also true of the closed source model providers as well.
>> Right. So last question for you Anastasia is what does Kim K3 I mean what does it mean for the businesses of anthropic and open AAI as you see it?
>> Well I said this uh online and I'll say it again. I think we're going to see a lot of companies that are distilling American companies distilling Chinese models now instead of the other way around. Uh I'm you know I'm sure open anthropic uh have their own way of training models but even be beyond those two uh you know the the Kimmy K3 is going to serve as a base model for many many US models um and we're going to see that happening now and if the trend continues that the Chinese models uh get better and better over time it may pose like a serious threat to the to the closed source business model because of all the drawbacks that the closed source uh providers have. Hey, are you expecting a comeback for for Google here? They've kind of fallen off the map here with the >> Absolutely. I actually have huge belief in the Google team. You should see those people. They're really smart. Their team is like excellent. Excellent. I mean, it's also true and open AI.
>> I mean, I just wondered if they >> No, they think they're good. I expect the comeback. Wait for it. Just wait for Gemini 4. I'll bet you it'll be, you know, >> we can't Well, hold on. We can't even get 3 point, you know, 35 is is delayed now according to Bloomberg. So, we got to get there first.
>> Yeah, that's right. That's right.
>> All right. Well, Anastasio, I want to thank you for coming on. That is Anastasio Angelopoulos, the co-founder and CEO of Arena here on TITP.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

WOW! Judge TURNS THE TABLES on Trump in His OWN $10B LAWSUIT!!!
MeidasTouch
197K views•2026-07-23

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Steam and Xbox Just Dropped The Hammer On PlayStation
OhNoItsAlexx
9K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23