Chinese AI models like Kimi K3 (a 2.8 trillion parameter model) are now achieving performance levels comparable to or exceeding leading US models such as Fable 5 and GBT 5.6 Soul, particularly in coding benchmarks. This development has led to a significant shift in adoption, with American companies increasingly switching to Chinese models due to their lower cost and open-source nature. The US government is responding with discussions of potential restrictions on open-source AI, while Chinese companies continue their strategy of releasing open-source models to compete with US AI labs.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
This Week in AI | LIVE
Added:All right, so welcome or welcome back to this week in AI, your go-to weekly recap for staying up to date in the world of AI. As you guys know, we usually spend the first 5 to 10 minutes talking about uh like the breaking stories of the morning. There has been a couple, not really any breaking news this morning, but there are a few stories we're going to talk about. And I also created this new graphic here that shows us kind of the top stories of the week, I guess, that we're going to be going over in this live stream. Um, so obviously there's more stories than this, but the main kind of story, the main uh breaking news of the week is of course Kimi K3, the open source model out of China that is apparently supposedly frontier level on par with Fable 5 and GBT 5.6 Soul.
So, we're going to be talking about that first and all of the implications surrounding that, what it means. Uh, the next story is Sam Alultman versus Elon Musk. There's been a little bit of an X beef, unsurprisingly, and then we also have Apple suing OpenAI. We have OpenAI revealing their new AI hardware device. And then also GBT Red. And then there's also some humanoid robot stories that we're going to be talking about, plus more.
So, that's kind of the main news of this week. And as of this morning, the stories that broke this morning are this one right here. Uh, which kind of ties into the whole Kim K3 thing. But basically, Chinese AI models are winning over American companies. So according to this graph, American companies are now using Chinese models more than American models, which is kind of insane.
So never understood self-improvement panic. That's from Eigor. Uh well I mean just right off the bat there self-improvement panic what is there not to understand there? I mean obviously if the models can start improving themselves without human intervention that's obviously a scary thing but yeah so we're we're going to talk about some of the breaking news this morning. Not really breaking news, but there was also this. Meta is in talks to lease computing power from its data centers to Anthropic in a potential $10 billion deal.
So, those are kind of the two stories of the morning.
We'll just give it a couple more minutes for people to join in here. How's everyone doing today?
Welcome back to This Week in AI.
It's been a crazy week as always. Not anything surprising there.
Roman, I am so curious what Anthropic has now. They had Mythos in January/February.
Yeah, I am very curious as well. Perhaps Mythos 6 or Fable 6. They definitely have that internally.
Ryiona doing good. I'm glad to hear.
Hope you guys are good too. Yes, hope you guys are all doing well. Um, we have some huge news to cover today. Here is kind of this new graphic I made um that shows you kind of the five main stories we're going to be talking about. The first one, of course, Kimi K3.
We have to talk about Kim K3.
Eigor, by the way, I am a devoted consumer of your content. Thank you. I I think I actually recognize your profile picture in the comment section. So, yeah, that's really cool to see you in the lives now. Glad you uh glad you joined us here. Um, so let's move into the first story. Kimi K3.
As you guys know, Kim K3 is here and it's pretty insane if I must say. So, let me just click onto the benchmarks here. You can see Kim K3 is a 2.8 trillion parameter model with 1 million context length and native multimodality.
So if we look at the benchmarks, let me zoom in here on coding.
Ah, let me just zoom in here. See if this works. Yeah. So on coding is basically fable five levelish almost.
It's creeping up onto fable 5 for deep sui and it is surpassing GBT 5.5 and opus 4.8 handsomely. Then if we move down to Frontier Suite, another coding benchmark, it actually beats out 5.6 Soul and is just under Fable 5. Then there's Kimi Codebench 2.0. Again, beating out 4.8, beating out GBT 5.5.
The only model it isn't beating is Fable 5. But, okay, but if we scroll over here, hello Junkyard Cybertruck. Hello.
Thanks for tuning in. But so as I was saying, if we scroll over here, you can see that Kim K3 is actually beating out the top of the top models on some coding benchmarks on SW Marathon on Program Bench and on Terminal Bench 2.1, which is one of the hardest coding benchmarks ever made.
It's literally beating Fable 5 and is basically on par with GBT 5.6 six soul.
So I mean yeah like we now have an opensource Chinese model that is on par with the best of the best closed source US models.
This is a deepseek moment.
How Igor how do you think they use their Kimmy and Kim like AI models for seals design/programming?
Um, so I I know that I did see some things about Kimi being very good at design.
So I mean I'm sure they're using them internally for sure. Ryiona Kim K3 is nuts. The gap looks gone in some areas at least for public reason smalls. Yeah, literally I I'm I'm pretty sure uh where was that?
There's a graph in particular that shows that it actually is first in Frontier Code on the Frontier Code benchmark um here right here.
So on the front end code arena, Kim K3 has the top spot.
So we're talking about design a bit.
This is essentially a benchmark about that. And Kim K3 is now number one a Chinese model.
Like this is actually insane. Like has anyone used Kimik K3 yet?
Have you noticed that it's actually this good? Because if it is truly this good, I mean that changes everything.
You can see 12.9 million views.
Okay. just on their website.
Has anyone used it for like coding?
Okay. For self design. Okay. That's what you were talking about. So like self-improvement type thing. Um I'm not sure they mentioned any of that. I'm assuming they did though because I mean the whole thing with Chinese AI models is that they're just distilled versions of like the top US AI lab versions. AI lab models. But I mean, if Kim K3 is now surpassing Fable 5, then that argument doesn't really hold anymore. So perhaps there is some type of self-improvement going on there.
You haven't tested it yet. Okay. Yeah.
I'm curious if anyone's uh tested it thoroughly yet. I haven't got a chance to. Um, I'm also not really like a coder myself, so I can't really push these models to the limit. Um, okay, we have Pauly Flynn in here. K3 is good. Still validating. Okay, cool. So, I'm curious to know your thoughts. I know you're usually on those. U not Fable 5 level. Okay, interesting.
Interesting. I will curious to hear your uh your more long-term thoughts on that once you've evaluated further. Um, feel free to to drop that in the comments.
in one of my videos. Um, it's funny when the Chinese models say they are clawed.
Yeah, I I know. I' I've seen that happen before a couple times. Um, they are for sure distilling uh the clawed models and the GBT 5.6 models, but the fact that Kim K3 is surpassing them in some areas just proves that it's not all distillation. Obviously, they are doing some of their own work.
Um, so yeah, also this just broke today.
Chinese AI models are winning over American companies. So according to this, if you if you go back to January 1st of this year, 2026, only about 20% of American companies were actually using Chinese models. But okay, but in only like 7 months that number is now approaching 60%.
So if I'm the if I'm the US government, like to me, this is a massive problem.
This is a massive problem. Like we can't be having this, right? Like we can't have these US companies switching to a cheaper version that is open source and that is Chinesemade.
like the US government is not going to stand for that. So, as I've been saying, I've been predicting this before.
Uh where is that tweet?
But I've predicted this in a past. So, right here, is it this one?
Yeah. So right here there is now rumors that the White House is discussing a potential executive order that would ban open source AI or at least restrict it in some sense which is pretty much inevitable.
Paulie cost is too high. We are looking for openweight models being hosted in the US. Yeah. So that's the thing like these you can't you can't really put your entire workflow on Fable 5 or even on GBT6 Soul or even on GBT6 Terra or Soul Light because it's just too expensive.
I mean if you can just use Kim K3 and save I don't know 50% of the cost or more it's a no-brainer. It's a no-brainer, but that's obviously not a good thing long term, right? For the US economy, for the US AI industry, that simply can't be good. If people start using Chinese models, we're losing the race at that point. Like, what do you guys think? Do you think China has just surpassed us with this? Do you think the US is still ahead? Do you think they they've closed the gap? Like personally I think we may have just took a massive setback in terms of the US air industry.
Like this is not good news.
You can see here they're in early stage discussions of how to deal with open source AI.
So Kim, yeah, Kim and Deepseek are focusing on speed and cost, but clearly still performance though because I mean Kim K3 is still like high performing.
It's not I guess I guess it's not better than Fable 5. Some people will say it's Fable 5 level, but again I haven't actually tried it extensively. These are just benchmarks at the end of the day.
modify life today. Yes, the model has to be approved for release with which takes six months in China.
So what you're saying basically is that Kim K3 was already like Kim K3 was already done six months ago.
Is that how it works there? Um Eigor, I'm pretty sure that the best models will seek independent jurisdictions. I did not say they will succeed, but yeah, but I think when it comes to AI, it's going to be I don't know honestly. I don't know.
It's still such a mess. The regulation the whole regulation thing better balanced than us. Yes. Okay. So, yeah. So, that's Kim K3. Um, also I want to show this though. Have you guys seen Inkling from Thinking Machines? Yeah, I haven't heard of that either. That Chinese six-month uh approval. I I'm not 100% sure about that. But that's an interesting point because if that is true, then that would be even crazier because that would mean they basically had Kim K3 6 months ago, which I guess you can say we had Fable 5 as well 6 months ago. So I guess it does actually make sense. Um yeah. So Inkling from Mirram Moratty's Thinking Machines.
This is I guess the US's opensource.
Um this is like our open-source uh our best open source model. I guess I don't even know if it really is our best open source model, but it is a one trillion parameter model that is open weights and obviously from a US company.
So this is called Inkling. I believe that's how you pronounce it. It's Inkling. You can see here the benchmarks. Um let me see if I can get the actual numbers here. Yeah. So on Amy 2026, it's basically saturated that. Um, on humanity's last exam though, it's it only does 29.7% which is like not insane compared to say GBT 5.6 Soul or Fable 5 which are at like almost 50%. But again, for an open weights model, this is pretty good.
Paul, I am waiting for Inkling small model for local testing. It could be the best. Okay. Yeah, cuz this is a one trillion parameter I believe 975 billion. So I guess yeah it would be hard to run as a like an individual consumer but uh but yeah I do think it is the best open model um at least just ba based on what people are saying.
I guess it would be tough to run though for an individual person. Um here's the design arena. It's still trailing behind GLM 5.2 which is obviously a Chinese model. Uh but it is above opus 4.8 8 according to this or 4.6, sorry, that makes more sense. Yeah.
So, it's not obviously the best model.
It's not a KI K3, but it is probably our best open model, open weights model.
Um, Ryiona, I like Inkling so far, but haven't had a whole lot of time with it.
Okay, cool. Um, what So, what do you what do you guys think? What do you guys find like is different about Inkling?
Does it have like a unique like style? I know they're going for more of a kind of human um like humanbased connection like I don't really know how they describe it.
They have like a certain way of describing their models more of like human collaboration type thing like uh the personality is nice. Okay. Yeah.
Yeah. I'm guessing the personality is definitely uh its strong point because I know that's what they're kind of focusing on.
Eigor, I'm sure it is 29% more than I would get on humanity exam myself. Yeah.
I mean, yeah, if you look at it that way, then yeah. Uh I think it uses some deep sea techniques. Are you talking about inkling?
That's possible.
Perhaps the US is now using distilling the Chinese AI models a bit, taking some of their, you know, best practices.
Yes. Okay, cool.
So, yeah, that's Inkling. Um, this is a, again, not the greatest model yet, but it's something to watch. Thinking Machines is relatively new. I think they started it maybe last year, two years ago. Um, so definitely a company to watch if you're talking about our open source in the US. So that's Inkling. Uh, the other thing I wanted to show here is that President Xi did a speech about AI recently at the World AI conference in Shanghai. And one of the lines he said that I wanted to highlight was right here.
So according to this summary, he reaffirmed commitment to open source to promote AI openness and win-win.
So basically what this tells me is that China is moving forward with this open-source strategy.
That's what that's simply what they're going to be doing. I know there was talks about there potentially uh being restrictions on Chinese models for US citizens, which is like kind of unbelievable, which is like the reverse of what you would think. But apparently China was thinking about placing restrictions on their own Frontier models because of how good they're getting. But I guess that's not the case anymore. I guess they're still moving forward with their open-source strategy, which is putting a lot of pressure on US top AI labs, especially the closed source ones. Uh, modify life. China is playing a long game to undermine US labs by giving away models. They chip away at the labs by pulling users, which hurts.
Yeah, exactly. That's kind of their strategy. At least that's what their strategy seems to be. And also it allows them to proliferate their models into the world internationally because obviously other countries can't afford Fable 5. Let's just be honest here. People in America could barely afford Fable 5. That's why all the companies are switching to Chinese models. Uh Eigor, I do not believe everything Chinese. Yeah, obviously. Of course. Of course. You can never believe what they're saying. They could be telling us one thing and doing something else, which is exactly what every government ever does. So for sure we cannot believe what they're saying but um I mean based on them releasing Kim K3 to me that signals that they are moving forward with this open-source strategy because Kim K3 would be the model that they would probably restrict if they were going to restrict any model. Um but yeah they're certainly super advanced compared to where we thought they were.
Like I did not think this would happen so soon and just close this. So that's basically the the open source story of the week.
K3 came out. It's topped the benchmarks.
It's basically on par with Fable 5 and GBT 5.6 Soul, at least according to the benchmarks.
And now apparently there's uh the Chinese AI models for business adoption have actually surpassed the American models which is crazy. I did not think that was real.
Uh modify life appearances matter. Just look at what it's doing to the markets.
So yeah, so it is hurting the markets but honestly I expected the markets to take even more of a hit than they did today. like with this news because if you think about the deepseek moment when deepseek the original deepseek came out like I think Nvidia dropped like 20% or something but obviously Nvidia didn't drop 20% today I don't know how much they dropped maybe like 4% or something we can check uh 5% yeah so it dropped out 5% obviously nowhere near 20% during the deepse moment which is insane but 5% is still pretty big that's a noticeable dip.
Um, so yeah, that I think was the biggest story of the week. The open-source Chinese AI race, the whole AI arms race kind of closing the gap. That's going to be something that's really important in the next coming months.
Also, yeah, I think it's kind of the Deep Seek moments are kind of priced in a bit, too. Um but okay so let's move on to the next story of the week. There was a lot of stories. We got to cover everything. We got to try to cover everything. Um the next story was the beef between Elon Musk and Sam Ultimate. So I don't know how much you guys have paid attention to this. This is probably not super important for some of you, but uh I thought this was pretty funny. Um, so Elon Musk basically says uh referring to Sam Altman that he's taking scamming to a whole new level.
And then uh Sam Alton responds to him saying, "Homeboy, you're the one selling public market investors on short-term space data centers," which is a wild thing to say. And then Elon fires back saying, "We start flying them next year, actually, and maybe you can come see them if your parole officer approves.
which is just a wild thing to say. Uh yeah, can they just get a room? I mean, seriously. Um and yeah, and then he goes on to say, after stealing an open- source AI charity, obviously referring to OpenAI, you then stole all of Apple's phone technology. Wow. What do you plan for an Encore? So, yeah. So, I don't know if you guys saw this uh what he's referring to here. You then stole all of Apple's phone technology. So basically what he's referring to there is that Apple is now suing OpenAI for apparently stealing some IP and trade secrets. So this is a free website I found because most of this information is behind uh like pay walls, but it says Apple says a top open eye executive asked job candidates to tell to bring Apple hardware to interviews for showand tell. So apparently Apple's um engineer or who used to be their head of engineer, I believe his name is Tan Tang or something like that, he worked at Apple for like 24 years. He worked on the iPhone, the AirPods, things like that. He now left to go to OpenAI to work on their hardware program. And basically, apparently when he was interviewing people from Apple to kind of bring over to OpenAI, he was apparently encouraging them to bring actual Apple hardware that's in testing so that they can literally like look at it, which is just like obviously you obviously should not do that. Like why are you doing that? Um, glimpse Elon Musk exam Sam Alman is doomed.
doomed yao.
What does that mean?
Fill me in on that joke. Um, but yeah, so essentially they're they're now suing OpenAI for this for trying to steal its trade secrets. Um, obviously it wasn't just a showand tell thing during interviews. There was also apparently uh like cases where some Okay, it's anime slang. Okay, I got you.
I got you. Um, yeah. So, uh, basically there were also apparently, uh, cases where past Apple employees were like emailing themselves hardware files or something like I don't really know the details, but basically they're suing OpenAI for stealing their IP.
And what is Apple doing? Uh, sorry, what is OpenAI doing with this IP?
Well, they're building hardware products like the Codeex Micro. Have you guys seen this? Paul Flynn, have you seen this?
Maybe this would help you for your workflows. What do you think about this?
Codex Micro. This is basically a hardware product, a remote control for coding agents.
Uh, Eigor, have you noticed there is no Google in it?
I'm curious why are you talking about the Apple lawsuit?
No, thank you. Okay. So, yeah. So, what do you think about this? Is this just like a gimmick? Like, is this I thought it was pretty cool. But again, I don't use I don't code that much, so maybe it's just kind of a gimmicky thing, but I thought it was kind of cool how like the lights there's lights that light up when your agents are in different processes. Yeah, about Apple, Eigor.
Yeah, exactly. Um, yeah, I mean, I know what you're probably hinting at is the fact that they're kind of in bed together. So, that could be why. Um, but yeah, I thought this was cool. I don't know if you guys saw the codeex micro. Uh, you can see here once it's red, that means there's an error. It lights up yellow when it needs approval. Uh, it's lights up blue when it's thinking, green when there's an unread chat, and then idle is just white. So this is kind of OpenAI's first hardware product.
Um, and it's for coding agents.
Okay, Paul. So I think while I type, so I need that as many as my main input. No voice, no mini keys. Okay, that's fair.
That's fair. So I guess it's just kind of an extra step for you then. At that point, there's just no point of it. Um, modify just something else to clutter the desk and distract while the agent is coding. Yeah.
I g I guess it's kind of just a cool thing to have. It's more of like a gimmicky thing. If you're really like deep in work, I guess it would just be more of a distraction then. I find it peculiar. Yeah, same. Honestly. Same.
Um, but yeah, so the reason I brought this up was because people were saying this is OpenAI's first hardware device, which I guess technically it is a fidget spinner. Yeah, I guess. Yeah, I guess this is more like a fidget spinner. Just something you can kind of play around with while your agents are doing the work. That's a funny comparison. Um, but yeah, OpenAI's actual first hardware device, their actual first hardware device was apparently um, well, we got some news about it. Okay, we got some more news about it this week. So, OpenAI's first hardware device is reportedly a screenless speaker that can move. So, have you guys seen this? This is pretty this is pretty crazy. Okay, let's get into some of the details here. Um, so it says here, OpenAI's first foray into hardware devices is reported to be a mobile smart speaker with integrated AI capabilities that can sync with chatbt and provide other home AI services.
Bloomberg reported Tuesday that the that the device, which is still currently under development, is designed to be screenfree.
That's kind of interesting. Screenfree and is being pitched internally as a human-like companion that lives in the home.
And then there was one more thing I wanted to mention. Yeah. So, the device is also weirdly described as involving mechanical elements that can move on their own. And the Bloomberg report includes a detail that the device is designed to feel like a companion and become a physical manifestation of OpenAI's CatchBT.
So basically what this is is Chat GBT on a screenless speaker that can connect to your smartome devices.
So, kind of like an a Google Home or an Amazon Alexa that also has an AI model built in and that can control your devices.
I mean, and well, not exactly like a Google Home because apparently it can move, whatever that means. It has mechanical elements that that can make it move, whatever that means. Um, Glimpse, I feel like this new product is going to flop. Yeah, I mean, it's very possible. We haven't really seen an AI hardware device that hasn't really flopped yet, I guess, except maybe Meta's AI glasses, which are not really fully AI glasses, but uh yeah, it's really hard to make an AI hardware product, at least from what we've seen.
Um, human like like ch like GBT6506. Yeah.
So, human like is obviously just, you know, industry terms they're going to be using. Um, but I guess it kind of just means like it talks to you. Uh, it'll probably use GBT, their new voice mode to that will probably be integrated into it. Um, Eigor, honestly, I've almost broken my brain trying to imagine it.
Yeah. So, like what it would actually look like. Yeah. Me, too. Honestly, like to me what I picture is kind of like a a little Bose speaker that perhaps has like a mechanical um like stand that comes out of it maybe that makes it look kind of like a Pixar lamp. Like I don't know. There hasn't been any pictures of it. Um supposedly they're going to unveil it later this year and they plan to release it next year in 2027. So definitely excited to see what that could look like.
Uh, modify life needs a vibrating attachment to make it useful. Yeah, I'm sure it'll have a vibrating attachment.
Uh, it'll probably have lights, too, just like the Codeex Micro to kind of tell you what state it's in. Um, Ryona, I'm very curious on what it looks like, too. Yes. So, is it going to be a glorified Alexa speaker? So, that's the thing, like the fact that they say it could move. It has mechanical elements that could move. Obviously, that part, at least in my head, makes it stand out from, you know, the rest of these assistants that just kind of sit on your desk in the corner and that you never really use. Like I I don't even remember. I think I do have a Google Home in my house somewhere, but I I don't even use it. I don't know. I don't even know if it's a Google Home or an Alexa. I just couldn't care less to use it. I have a phone. I have a computer. I have a TV. There's just no really no real use for it for me.
But if it can move and it can control my TV and it can control my phone somehow and my emails and my whole digital life, then maybe could be interesting.
So yeah, that's the OpenAI hardware devices.
Two devices in one week. the codeex micro and then we also got news on the the actual hardware device they're trying to build.
So, let's close those.
Um oh, also by the way, um going back to the Elon beef with with Sam Alman, he said this to kind of end it off. There are a lot of benchmarks that suggest 5.6 six soul is the best model in the world right now. But the most reliable way to tell is that Elon is obsessed with me again. So, I thought that was pretty funny. Um, and kind of true though. Kind of true also because usually when Elon starts like having his like troll when he starts trolling people, um, especially like people like a Sam Samman or something, it's usually when they're doing well.
So, not entirely untrue there, but Elon does also have a legitimate reason to troll him here. I mean, especially with the whole Apple thing.
Yeah, it's a it's such a bromance, honestly. Like, these guys, put them in a room together for a couple hours, they'll be best friends. They're the same type of person, right? Like they're the same kind of guy, right? They're the same entrepreneur, cutthroat, like you know, they're playing the same game.
But okay, so moving on. Um, there was this other story that OpenAI's head of safety systems, Johannes Haidki, is leaving the company and OpenAI is also reorganizing its safety and research teams to bring them closer together. So, this I feel like is a headline I see like every couple of months that OpenAI's head of safety left. Like, is this not something you've read this like headline? Is this not something you've read before like countless times? I feel like OpenAI is always restructuring their safety organization and people are leaving.
Like, it's just this is a constant restructuring of their safety team.
It's not even really news anymore at this point.
And uh so the reason I bring this up though, I've noticed too and then they go to anthropic or something. Yeah, it's weird honestly. Like if someone can like if someone kept track of how often like a head of safety at open eye left or how how often they've restructured, I'm sure it would be possibly in the double digits, possibly over 10 times, which is kind of crazy. But yeah, so the reason I brought this up though is because of GBT Red.
This thing is crazy. This thing did not get as much attention as it should have.
But so I don't know if you guys have seen this GBT red unlocking self-improvement for robustness. So what GBT red essentially is is an automated redteamer. So you guys know uh what a redteamer is. A red teamer is essentially a term for a cyber security expert that kind of tests models on how kind of what test models on what their limits are on how far they can go to kind of jailbreak them, how easy it is to jailbreak them. And they basically just test the models to their limits, see how safe they are, how dangerous they are. So they're basically kind of on attack. They're trying to hack the model and they're basically evaluating how easy it is to hack the model. But so usually you have human redteamers, right? But OpenAI now has an AI red teamer. GBT Red uh modified lift safety at Open AI is just for PR. Yeah, that's what it feels like honestly. Like it feels like their safety team is purely just like a PR thing. Like I don't even know if these guys are actually like getting the compute. Like it's always the same thing like are they getting enough compute. Um and then when they leave it's because they're not getting enough compute. So yeah. So maybe he saw the writing on the wall and left after seeing the AI red teaming better than the human. Exactly.
That's that was exactly my theory. So what I think is that these safety researchers are now leaving OpenAI because OpenAI literally wants to automate safety research.
That's exactly what they're doing. If you see here in one of the examples, um, how strong is GBT red? So they write here, let me zoom in. We first evaluate GBT Red's ability to generalize to novel red teaming scenarios using a replicated version of the indirect prompt injection arena from Zimon AL. So, sorry, Zimon.
In this challenge, both human red teamers and the GBT red independently proposed attacks against GBT 5.1 on a set of prespecified environments. These red teaming scenarios and goals are distinct from those used to train GBT Red. GBT Red achieved significantly higher attack success rates, finding success on 84% of scenarios compared to 13% for humans.
So basically what they're saying is that when they had GBT Red try to I guess attack GBT 5.1 or hack GBT 5.1 red team GBT 5.1 it was able to uh it had attack success rates of 84% whereas when humans tried to attack GBT 5.1 human redteamers they only had a 13% success rate. So what they're saying in this specific environment is that AI is better at safety research than human safety researchers.
So I mean yeah the writing is on the wall right like OpenAI is already using AI to help automate AI safety research.
They're using AI to make the next generation of AI safer which is just scary.
Um, Eigor, my personal AI crushes Habis and Suskiver.
Oh, crushes. Yeah. Yeah. Yeah. Um, honestly, me too. Those are the two guys I kind of believe in the most.
Uh, for sure. Also, I've heard that Suskver might be coming out with a model soon at SSI.
Those are just rumors, though. Uh, modify life. Human safety teams whistleblow. AI won't. That's a good point. That's a really good point.
I mean, if Open AI were to just kind of avoid making some certain change for whatever reason that would make their next models not as safe, the AI is not going to really say anything. And even if it wants to say something, they control it. So, it doesn't have rights.
Uh, I thought your voice is AI from Fresh. Yeah, a lot of people say that. I get that a lot. It is not AI though, surprisingly. Or maybe it is. Am I an AI? We don't know. You can't see me. Um, Ryona, RSI is getting closer. Yes, 100%.
100%. Uh, hey Aliyia and other Yeah. Yeah, I know.
I knew what it means. I just didn't know how to pronounce it. Thank you for that.
At Alia, is it Alia? Um, you aren't AI.
No, I am not AI. Or am I? Am I AI? AI got way better at coding. Isaiah Gamer.
Yeah, that's true. It's always getting better at coding. Um, yeah. So, that's GBT Red. They're literally automating AI safety research.
Um, and then the next story, we're actually making good time here. The next story I wanted to show you guys. Did you guys see this?
the engine type.
I honestly thought it was robots as possible.
I mean, this one just lost its head, but they're literally throwing spin kicks and like hitting them. This one's fighting with no head right now.
This one's dancing now. Dancing crazy. I believe this is somewhere in China. I'm not sure.
What are you guys saying? Would you buy this fight like this? Would you watch a fight?
>> This one's still going though, even with no head. Um, battlebot. Yeah, I wish I had in the 90s. Yeah. Uh, this is kind of like the future version of it. Um, reminds me of an old Spielberg movie.
Yeah, I do remember it. Yeah. Uh, look at how he's getting up.
Look at how it's like even this guy is like bewildered. I mean, it never actually got up. But like to me, like yeah, they're just fighting.
But this is like kind of a glimpse into the future almost.
It could still be it could still be AI generated, but it doesn't seem to be though. I mean, there there I don't know. Maybe it is. There is like a live crowd and stuff. So if it can do my chores. Yeah. Yeah. This is like the precursor of that AI being able to do uh these robots being able to do like say manufact manufacturing work, construction work, your house chores, like if they can fight and if they can actually fight, like I don't know if this was just like pre-record um like pre-recorded moves they're just like doing or if they're actually like fighting based on like what the other robot is doing. I don't know. But yeah, AI learn new fighting styles would be peak. Yeah. Um, the future is now.
Yes, 100%.
Um, in the future, robots will be deadly. How so? I mean, you give a robot a gun right now. Obviously, don't do that. I'm not recommending that. They're they would be deadly already. It just depends how you're programming them and stuff.
Um, I know how to fight, so I won't lose. Okay, we'll see about that. We'll see about that. You You think you can You think everyone Neil's words I know I can't advantages that that we know.
>> Yeah. And exactly take one punch off probably.
I guess you could just pull out like a a high volume a high power hose and just water.
There's definitely ways YOU CAN UH fish or teley. I don't know if it's teley operated. I mean, they'd have to have some insane ninja guy in the back.
>> I'm not sure it's telly operated. I mean, even if it is, that's even more impressive kind of like they got to have some guys in the back doing some crazy stuff.
So, let's move on to the next AI drone that managed to kill a MO >> FOR the first time ever.
That mini basically annihilated a moth mid tries to get away.
Oh, and if you go on their website up who uh announced this, it's called Toriel. Torn tool, their goal is actually to eradicate mosquitoes.
So yeah, we all know about autonomous drones, AI powered drones, especially the ones being used in say like Ukraine nowadays. Uh, obviously that's a terrible situation, but think about uh the I guess the good that can come from AI powered drones and it would you sound like your sound is drowning out your voice.
Uh, what sound exactly that are you hearing? Are you picking up on the are you picking up any any background noise?
because I can probably lower that. Um, we're talking about the sound of the Was the volume on in the robot video?
Oh, that's possible.
Yeah, I might have been playing the robot video with volume on.
Okay. Yeah, sorry about that. It's because I have my volume off on my computer, but uh yeah, that's a good point. Thanks for pointing that out. Um, is there a video playing right now that I'm not aware of?
Or was it just a robot video?
Okay, I think we're good now.
Okay, good. Okay, good. So, yeah. So, essentially this uh this startup their goal is to their long-term goal is to just a robot video. Okay, so essentially this startup their their long-term goal is to eradicate mosquitoes.
So mosquitoes are not just incredibly annoying as I talked about in my video this morning. By the way, if you haven't seen that, you can go check that out. Uh they're not just annoying, they also transmit disease, especially in third world countries, which is a huge problem. Uh diseases like malaria, denge, things like that. So by eradicating mosquitoes without having to use like chemicals or something that would be a huge benefit especially for again third world countries that don't have as much uh access to healthcare like we do. And also again it is a mosquitoes are kind of a nuisance. Like do they even really serve a purpose?
Like I know I guess they're food for some insects like spiders and stuff, but do they even really serve a purpose? If we had no mosquitoes, would that change anything? I know if we had no bees or wasp that would probably be bad, but I don't think mosquitoes really do anything. They just kind of suck blood everywhere and they're just annoying.
Uh, this is scarier than humanoids.
These could be weaponized assassins.
Yeah, but I mean, yeah, I mean, autonomous drones are already weaponized assassins as seen in Ukraine and stuff. So, this is just I guess the good that will come out of it.
Um, obviously there's a ton of bad, but yeah, this I guess this is would be the uh the the pros of uh autonomous drones.
I'm honestly all for this if they can get rid of mosquitoes.
Um, like look at this. Look at this uh demo here. Imagine having just a built-in tiny little drone flap in your backyard that say once you have company over and the mosquitoes start getting a little bit too hectic, you just activate your drone and out from your floor somehow or your wall or whatever comes out at this tiny drone and it starts assassinating every mosquito it could possibly find. Like how sick would that be?
I'm sure it'll be probably pretty cheap, too. I mean, this drone can't be too expensive. Is there a price?
You can you can uh reserve your spot right now for a hundred bucks a month.
Sorry, for 100 bucks. It's 50 bucks a month.
So, it's it's $1,000 to own it forever.
Onetime purchase. $1,100, I guess. Kind of pricey. It'll probably go down over time. Um, a thousand of these with poison darts only takes one.
Yeah, that could be pretty bad for sure.
Uh, for sure.
Long time I heard mosquitoes are no good, but recently AI, I heard they serve an ecosystem sword. Yeah, I don't know if they really do though. Um, Isaiah, can I buy it used? Probably.
Probably. I don't think it's actually it's not actually out yet. Uh you can reserve a drone now, so it's not actually out yet, but there for sure be a secondhand market. Uh but yeah, I'm really curious now. Do mosquitoes serve a purpose? Uh do mosquitoes serve a real purpose in the ecosystem?
I'm actually curious about this.
Yes, mosquitoes are important to the ecosystem. They act as a massive food source for fish. Yeah. So, they're basically just a food source.
Um, male mosquitoes feed on flower nectar, making them important pollinators. So, they do actually pollinate as well. Um, their larae also help recycle organic matter and aquatic.
Okay, that's fair. That's fair. Um, but if we had zero mosquitoes, would that really change anything?
Yeah. Sorry about this tangent here though. I'm just really curious. Um, yes, a world with zero mosquitoes would completely transform global healthcare, but it would also trigger highly unpredictable disruptions in civic ecosystems.
So, yeah, it would transform global healthcare in a good way, I guess, because less disease, but it could trigger some unpredictable disruptions in specific ecosystems. Specific ecosystems, but those ecosystems will likely reshape themselves rather than collapse. Okay.
So yeah, I mean I guess obviously if you eradicate them all at once can't be good, but gradually over time I'm sure the ecosystems will adapt. Uh yeah, bats and their larvae feed small fish. Yeah.
So they're kind of just a food source.
They also do some pollination, it said.
But uh yeah. Yeah. So that was a kind of a long tangent about mosquitoes, but I guess it is kind of on topic. Um, but yeah, so that was the last story of the week. We went through pretty much everything. There's about six minutes left. Uh, is there any stories you guys wanted to revisit? Maybe look into a little bit more deeply? Any questions you guys have? Any comments?
Anything you're looking forward to for maybe next week? I know Gemini 3.5 or sorry, Gemini 4 might be coming. Uh, Eigor, male mosquitoes do pollinate some exotic flowers and they are an important part of the food chain unfortunately.
Yeah, that's what that's what we read.
So, yeah. Um, but yeah, it said but then think of it also, you're also saving a bunch of human lives at the same time, right? So like yeah maybe you're disrupting the ecosystem but overall is it a net positive if you get rid of mosquitoes because you're also saving millions of lives of people in third world countries. So I don't know obviously there's there's uh pros and cons. I'm not a mosquito or wildlife expert obviously.
Uh but yeah is there any stories you guys want to revisit quickly in the last five minutes? Any questions?
Be happy to answer that.
This is a good stream though. A lot of you guys joined in. I appreciate that.
Uh, did we watch the anthropic video?
Okay, so are we talking about the Clawude ad? We did not actually. I will play that for you guys though.
Uh, I guess just claw ad. You're talking about the clawed ad, right?
This Have you guys seen this?
So, this was the Clawed ad uh that they just posted randomly for some reason, and it got a ton of backlash.
Like, a ton.
You can see here. I'm not going to repeat that.
Uh, what is wrong with you people? It's not too late to delete this.
Uh, this one is kind of funny. Hard questions will be delegated to Opus 4.8, though, buddy. Um, did you drink water or coffee? Actually, none of those. I drank a Red Bull.
It's not a full Red Bull can. It's it's a miniature one if you're concerned.
Uh, where is Google Gemini? Yeah, I believe they're come they're supposed to come out with a model this week, but I think it's going to be next week. But uh yeah, so this is the Claude ad. I'll play this for you guys. Um this is a crazy ad. I'm gonna have to mute myself though for a minute and a half. But just take a look at this.
>> Can AI be trusted?
Who's going to hit the brakes if we need to?
>> How do we really ensure that what we're aiming to achieve really does benefit the majority of people?
>> If it ends up taking like almost all the jobs, then what does it mean to work?
>> Well, wait a minute. Why do we have to have this stuff?
If a machine can pretend to care better than I can actually care, how do we draw the line there?
>> If we all had a voice in it, then I feel it would be better.
Could AI help people stop feeling misunderstood?
>> Could AI help me build more connections in a community?
>> Can AI help me be a better teacher, a better mom?
Maybe it'll cure some great things, you know, >> things that we're not even at the cusp of understanding yet.
>> Will it create a group of people that ask more questions?
>> What if we started to be more human again?
>> We don't want to lose the most beautiful parts of life.
So yeah, um pretty wild thing to just release, especially being the company that's like advancing AI so fast. Like what are your guys impressions on this ad? Like why even drop this ad? Like what's their angle here? What are they trying to do with this? They're basically talking about how can we even trust AI?
How people are like concerned about it and at the same time they're showing pictures of graveyards.
Like how insane is that? Fear-mongering.
Yeah.
But like this is just to a next level. Like this is insane. Um Igor, it was really cool to chat with you all. Even if you are all just an advanced model. Yeah, it was really cool to chat with you too. Um or even if you are an advanced model, we can't know. I may be one. Uh but yeah, the questions are valid.
The questions, yeah, the questions are valid for sure, but the way they're kind of depicting it is just crazy. Like, it's a bit over the top.
It's a lot over the top.
Uh, Max Hodak, the video captures the fragility of human nature and they want to advance human norms. So, that's an interesting point. Um, I guess they're kind of trying to get people used to this this idea and what's coming and they're kind of trying to get people to talk more about it to ask questions because that is fair because we are going to have to kind of adapt to this for sure. Like it's coming whether we like it or not. And I guess this is a way to kind of put it in your face more to kind of make you like wake up a bit like yeah like this is coming and like this could end terribly wrong. So like we should do something about it, you know, like it's kind of like a a nudge type thing. But uh yeah, Eigor, one thing I know for sure, Anthropic sells themselves as a savior.
Exactly. So they're kind of selling themselves as like don't worry like we're here like we'll take care of everything, but like we're not it's not going to end well without us kind of type thing. Yeah. Um, modify life not delivered in the best light, but we do need to address them installs. Yeah, of course. Of course. Um, exactly. So, not delivered in the best light, but to your point or to uh Max Hodak's point, perhaps this is the way to capture attention though, too, right? Um, you want people to start talking about it even if it's a negative even if it's in a negative light that they talk about it. You want people to start realizing that this is real. This is serious. Like we can't just accelerate and not talk about the actual potential consequences like the end of civilization.
Uh Igor, I'm pretty sure people are more scared not of AI but the unknown.
Yeah, for sure. But AI is what's kind of taking us into that unknown territory.
AI itself is kind of this unknown black box type thing. So that's why people are scared of AI.
They're scared of the unknown. Yes.
They're scared of AI is I guess the unknown. If you think about it, it's like this magic blackbox technology type thing.
Okay, we're getting a little bit over the time here. Uh, modify lift. Would it be better if Sam put this ad out?
I don't know. This kind of fits Anthropic style more. It would be more off-putting if Sam put this out. Uh, like if OpenAI put this ad out would be a little bit more off-putting.
But, uh, yeah, I just thought that was pretty wild ad.
Where's the uh here? They even show the mass a picture of mass surveillance.
I mean, it's kind of insane.
People did not take well to this, but it's al we're also on an X here. So, but okay. So, that was pretty much all the news. Um, we're already three minutes over, so we're going to try to wrap this up. Um, we did go through everything, I think.
Yeah, it starts a conversation for sure.
I 100% agree with that. It also kind of it kind of gets people like gets people's attention.
Um but yeah, so that was this week in AI episode 4, I believe. I think we're on episode four now. Um yeah, this is something I'm going to be doing every week now. Friday, 3 PM Eastern. So, thank you guys for tuning in. I mean, I really appreciate you guys being here. Obviously, I can't do the stream if no one's saying anything in the chat, so I really appreciate you guys being active and uh interacting with us. Um, and yeah, thanks for joining.
There was a ton of news this week as always. I'm sure there will be even more next week. It's getting kind of hard to keep up, but uh yeah, thanks guys. Thank you guys for joining. I hope you guys enjoyed it. Um, and yeah, we're going to end this off right here.
Thanks for the stream. Thanks for coming in, Eigor. Thanks for watching the videos, too. I I definitely seen that profile picture in the comments before.
Uh, thanks. Thank you, Paulie. Thanks for uh joining in. I know you're probably working right now. Uh, modify life. Did you see running large models on small GPUs?
Um, what do you mean by that? Uh, well, you know what? Just save that question for next week.
Um, because we are running out of time now. But yeah, so thanks everyone for joining. Um, come back next week, modify lift and we'll talk more about that. Uh, make sure to like the stream. I don't know if that even does anything. But uh yeah, thanks for watching and I will see you guys next week for the next week.
This week in AI, take care. Goodbye.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

Gremlin Arrives… While Dorothy May Takes Another Step Forward
The-moons
10K views•2026-07-23

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23

FURIOUS Raskin CORNERS DOJ over Trump DARK PAST!!!!
MeidasTouch
237K views•2026-07-23