Sentdex correctly identifies that for real-world productivity, low latency is a form of intelligence in its own right. The shift toward high-speed local inference marks the end of the "bigger is always better" era in practical AI deployment.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
You Can Just Download More Tokens/Sec
Added:What is going on everybody? Welcome to another Frontier AI at home video. Uh, since the last video, so much has happened. It's still less than two weeks. We haven't talked about GBD56 Soul. Kim K3 came out. Thinking machines is releasing openweight models out of nowhere. Like these are ex open AAI people. Uh, and then uh just before I turn on the camera, I saw Quen 38 is releasing. So, both Quen 38 and Kim K3 have announced and they're available, but the the weights are not released yet. And any day now, I suspect our friends at DeepSeek will be releasing another model. But the idea here is that you can just download more tokens per second and get faster models. And I I don't think in the previous video you even got to see just how fast these models are because I was using the same uh machine to record as I was running the AI. That has changed. More on that later. Uh, but I want to show you the speed because this is uh what I am going to reference a lot through this video because I'm going to argue vehemently that speed obviously you need both. You need speed and intelligence. Let me go ahead and just run it.
As you can see it be zooming. Okay, this 302 tokens per second. Obviously, we have a very small context, but that really doesn't matter. But time to first token also really matters. It's these two. It's like all of this about speed.
And um so I think what I'm going to do is so you can see how fast that was.
Let's run one more thing. So these are just like it's a it's a prompt I run. I use this to benchmark speed. Um but I would do that like headless. But I just want to show you visually what the speed looks like. So I'm just going to keep copying. So it's just asking to do codegen. And codegen tends to be the faster version. Like generally you can decode almost twice as fast as like regular uh output. So like regular text output would be maybe a little slower.
322.6 tokens per second. Oh my god, it's so fast. It's so addicting. Um and again the tokens per second super quick. We're we're we're still very small context. So now what I'm going to do is just in the background I'm going to ask um Minion to look around on itself and um read all the files into context. Hopefully we'll get a nice big context and I can show you another point that I want to make with Deepseek V4 Flash and with DSpark.
And the idea here is um I might I think we'll be okay actually. So already you can also tell like it runs like all these tool calls, reads all these files like everything is so fast because every tool call itself is a um it h you have to interface with the model. So everything that you do and every like all the in order to know what parameters you're going to put into a tool call, you have to have context. So you're still passing everything literally through the model. So in this case, you know, anytime that we're kind of waiting for a moment, it's actually we are we're just simply waiting for the actual tool call like the actual command itself to take place. Okay, so we're at 90K. This is perfect. [laughter] I can't believe we we filled it so fast. Um I'm trying to think of a good a good question to give you. So, so what I want to say is so two apparently two that might be like the tool call. We'll see. We'll see. So 90K um I'm trying to think. Um I think what we'll say is rate I don't want I have no idea what it's going to say.
Rate this repo uh 1 to 10. Um yeah.
Okay. So what I want to say is so like on GLM52 local uh the the time to first token was really slow because your prefill speed so how fast does the model load onto um load context into memory basically before it starts generating new tokens. So generally your prefill is much faster than your generation speed like way way faster but different models have different prefill speeds and this makes a huge difference. So like GLM52 uh local at like 4bit for me um the prefill speed is like on a good day it was like 1100 tokens per second. Well we have a context now of 90,000 tokens. So 90,000 tokens at 1100 tokens per second is like I don't know 80 tokens per second or I'm sorry 80 seconds [laughter] um just to fill up the content. So, so when you send a a request, like you send a new message and you want to get a response from your AI and do inference, it's going to take 80 seconds just to start seeing output at all. And this is why like when you're talking with like claude or even codeex like if you have huge co uh context, it's slower. It's not going to be as slow as GLM52 local um but it's slower. So now what I'm going to say rate this repo 1 to 10. We have 90k context. So again, it would take 80 seconds with GLM52. And in this case, it took I don't seem like less than a second. We'll see what it says at the end here. [snorts] Um, yeah, less than a second. 724 milliseconds. And the reason for that is the prefill speed on DeepC V4 Flash. So again, GLM52, 1100.
So like what's say 1,000 DeepC V4 Flash, it varies depending on the size of your context, what's in the context, but we're talking like 100,000 tokens per second. So again, it's like the prefill speed almost doesn't even matter. Like I didn't know until I wanted to tell you what it was. I had to like look because I have never even looked what is my prefill speed on DC4 flash because it didn't matter, you know? Uh but GLM52 is baked into my brain because it matters. And I spent a lot of time trying to think, man, how can I make the prefill faster because this is painful. But like when I was using Miniax M27, somebody asked me once, what was the prefill speed? And I was like, uh, I have no idea because it's irrelevant. you just don't feel it anymore, you know, and so it just it really didn't matter. But in this case, like even if you went out to a million context, it actually goes faster. So So the bigger the context, the more your prefill speed will be. I consistently see, unfortunately, we did see like allegedly a two I I never saw it pause for 2 seconds, but um but allegedly that's what it was. Uh I almost never see a prefill speed higher than like 1 second as I'm using this model. And man, it just it makes a huge difference. And so like repeatedly I'm going to hammer this home in this video that I think like intelligence matters. Like you want a smart model, but at some point you you are trading intelligence for speed as we make these models bigger and bigger and bigger and bigger. [laughter] And I'm going to argue maybe it's not always the best to be the fastest.
Maybe it's not a it's it's also like you don't you don't want to maximize fast and be stupid, but you also don't necessarily want to maximize intelligence and be super slow. Like it's useless. And uh so so we'll we'll talk more on that as the video progresses. But uh so yeah, this is this is with with that D-Spark. And so DS-park is Deepseek's latest iteration on top of like all the advancements of speculative decoding. And already DeepSeek V4 flash was using MTP and they were actually using MTP layers in the model. So uh the the concept of speculative decoding the whole point of it is every like like we we have said multiple times also in this little series these models are just next token predictors.
They're not actually you know super intelligence or anything. They just all they do is predict the next token. Well, the problem is that's a little inefficient sometimes, especially if you have a really big model. And even Deep Seek V4 Flash, even though that's theoretically small for its performance, it is still a pretty big model. And so every time you want to predict a token, you have to run it through that whole model and it's very slow. And so the idea of speculative decoding is it started as like small draft models that like like a completely separate model that you were running and that model would predict the next x amount of tokens. or end tokens and it would be like two three four maybe five tokens or something like this and then you'd use the big model to verify those tokens in parallel and this was quite a bit faster and like sometimes it may like you I don't think I ever really saw anything get actually double the speed with just a basic draft model. I'm trying to think if I ever saw that. I I think you you tend to see though like maybe a 50 a 30% 50% speed increase though. It was still significant. it was enough that you would use it. And then models started doing what Deep Seek V4 Flash was doing and have like actual MTP like layers in the model itself and then doing some more fancy things. And then DSpark is even more on top like DSpark is doing a lot of so like it's essentially just a bunch of software stuff and it's also attaching like a sort of uh MTP module or so. I don't even sure if they want to call it MTP or I'm not sure what they're what they're calling it internally. Uh let's click let's I think I have the paper here. Um, I'm trying to see if I can find or maybe it's here somewhere.
They they mention a module. Yeah, they just call it spec an addition additional speculative decoding module. Um, so anyways, and it's at certain layers on I forget it's like 40 and something like that. Anyway, whatever. Uh, the bottom line is it's faster and on top of already the speed gains we were seeing with just the MTP layers, we this is another at least on my hardware which could be improved is 50% faster. So again, I I've probably said this a million times. So like here, we're seeing a little over 300 tokens per second. I believe with um with with DSpark on top of PCIe 4 and if we were at TP4, that's the other thing. I'm actually only on TP2 here. I'm using two RTX Pro 6000s only, and I'm getting 300 tokens for It's kind of It's kind of insane. Uh, if I was on four of the RTX Pro 6000s and running x16 PCIe four or five, um, I should be seeing more like 400 tokens per second. Wild. So, yeah, it's it's very fast and all you have to do is just download it [laughter] and that's it. Um, and, uh, I do want to show the the, uh, performance benchmarks because this was interesting. So actually DeepSeek V4 Flash with DSpark scored better than GLM 52 on open router and this was FP8. I don't yet want to put on blast who the provider was here but I am very strongly believing that this provider was at least at that time even though it said FP8 I think they were providing FP4 that it just the the result lines up with being an FP4 model.
So, um, yeah, more on that maybe, maybe maybe one day. I don't want to call anybody out and I don't want to like get it wrong, but I just I know what the number is and I know what the number should be. [laughter] It just it's not right. So, um, anyway, uh, I would like to maybe like benchmark multiple of these models or multiple of the providers on Open Router. And I wish there was a way on like Open Router that you could like more quickly detect if a because it's so easy for a provider to just not do that, right? It's the same reason why like like I have a list of reasons of like why local. It's like the model will never change. It will always run at the precision that you expect it to so that that behavior won't ever change on you. You protecting your IP.
It's actually faster. And I'm going to really try to drive that point home because you can't even you can't achieve the speeds that I just I'm showing you today. You the only way you can achieve those speeds for yourself is if you are running local because all the API providers basically target the 50 to 80 tokens per second range. Every now and then when I'm running GLM52 on together like their API I do see sometimes in the 200 pluses but the the thing that's very frustrating is sometimes it's 200 tokens a second sometimes it's five [laughter] because it just depends on what's the load on their API and obviously if you're running on prem and you're a big company you've got 500 users and you've got a server rack you're going to experience a very similar thing it's not it's not the the the performance is going to fluctuate so um I don't want to I don't want to slam together too much here uh for that either. But anyway, um I would like to test more of the providers on Open Router. Somebody from Open Router actually told me like if they would just give me some credits and I could do it. But like I said, I I feel like you want to be really careful if if you're going to start calling out providers, but I'm pretty sure Terminal Bench V2, I know what it should be and that score was not right. So anyways, um take that with what you want. But still, Deepsec V4 Flash versus also Deepsec V4 Flash. Both of those are run local on my machine. We've got an eight task difference and 65 versus 56. Boy, that's a big error bar. Uh I don't think there's any reason why cuz this is running MTP2 with the built-in and in that version of spec some correct me down below if you actually know what you're talking about. I believe like both of these versions of speculative decoding are like theoretically lossless. Like there is no reason why there should be a degradation of performance. So these two models should have scored identically and they didn't.
And it's probably just error bars. But boy, that's a big error bar. Like eight tasks, like my goodness, because like well eight tasks here would put GLM52 to the mid70s where it belongs. So it's h it's it's it's interesting. Um and that's exactly what I mean is like it costs on open router like 607 $80 to do a benchmark of GLM52 for example. And I think if you really wanted to be fair, you'd want to at least do like three runs per provider. And that gets expensive really quick.
But that's how you'd have to do it, I think. Um, if you really wanted to confirm somebody was like running what they say they're running. Anyway, moving on. Um, I want to show the tokens. So, this this table, this is like the first time I'm really looking at this table since I had Deepseek V4 Flash write this table. So, I'm going to try to make sure nothing is incorrect here, but I believe this is all correct. It is important to take note that this is TP2. So when I say TP2, not only are we running like VLM at TP2, it means I'm actually only on two RTX Pro 6000s. So in this case, this was on four RTX Pro 6000s. This is on two. And if you add two more, it this number alone will get faster. And these are like kind of peak numbers. So 185 versus 330. And then as we continue to scale out to eight concurrent, this is per uh per concurrent um connection. I don't know why I'm bling on this kind of term, but anyways, per uh per chat or whatever. Um and then here we have the aggregate. So as we continue to span out and at what one thing that's pretty cool is like a single a single uh conversation concurrency, whatever. Um I'm going to keep tripping on this, I guess. Uh is like way way way faster. But as we continue to span out, I think I mean we're asking a lot of these two uh RTX Pro 6000s. So we do finally fall below at eight concurrent. But I mean again, we're talking two RTX Pro 6000s serving eight concurrent um usages at basically 150 tokens per second. That's crazy. And it would get faster if we were on like PCIe 5 for example. So yeah. So now let's talk about um my updated hardware.
So this is the current situation. Now every time I show hardware and start talking about hardware, people get a little angry about money. And again, you can do a lot of this stuff without spending very much money. You like Deep Seek V4 Flash a little tough to run. You can run it on other hardware, but I do want to share how it's done because you are spending a lot of money if you are going this route and I want to help other people save money if you are spending this much of money and share what I am learning along the way. So uh so this is in the Puget build which is like a 4 and a half old a 4 and 1/2 year old uh computer and but it has so it's like has like seven PCIe four slots.
Some are x16, some are x4 x8. Um, and then also the crown jewel actually is probably the RAM at this point. So, this is a terabyte of ECC RAM. It's DDR4 though, so it's like old older RAM, but it's very nice RAM. Um, so we could use some of that potentially with some of these new models that just recently came out. We will talk about those hopefully soon. Um, so as you can see, there are actually three cards here, but when we are running DeepSeek, all the things I just showed you, you can't actually run Deepseek with TP3. Like there there's too many. I believe it's the attention heads like so you have to have the attention heads divide by whatever number. So like if you usually they divide by like two so you can always do like two four and so on. Um but Miniax M uh M27 I don't think M3 actually did have TP3. I can't remember but Miniax M27 had 48 attention heads. So you could do TP3 because it divided by three. So it's actually kind of a cool a cool model but it's pretty rare. Usually you have to do it in like in like times two increments. Um so anyways uh yeah in this case we are actually only running on these two and these are two cards connected to x16 PCIe 4. And um moral of the story is risers suck. I hate them.
They're terrible. Uh they you shouldn't even be allowed to sell risers.
Everybody that sells risers needs to go to jail. Uh the the problem was I I think there's many problems like you have the signal integrity over a distance. They're not like powered so it doesn't do a good job pushing that. I mean they're kind of powered but not to go that distance. So you lose signal and then um also you know this is where the PCIe slot is. Well this is also where you got a fan on the bottom and it shoots air up. So these are like these are uh what do you call it? Heat sink and then it's shooting the air up. Well, if you have a a PCIe cable going in to one of the PCIe slots, like here, like on my old build, there's a little more space between the cards, but you still had a a cable going in through there, well, guess it's just it's just getting baked by the card that's under it. And so, I think I think what what really did me in was I actually had pretty good stability on those PCI or on those risers at PCI PC PCIe3. Wow, that was really tough. Um, I had good s I had a good stability there and then I did like a week of benchmarking and like running really hot like overnight and I think I just like baked those cables or something and I have more but it's like why would I do this? Like I can't um like I can't even run on my main computer. I can't run for more than like 10 or 15 minutes when I was using those risers without my computer like the PC the cards are just falling off the bus and the the computer just freezes or cra like crashes. I got to restart every single time. It's just it's obnoxious.
So the question though is well there's seven PCIe slots man and then also like even though I could fit three cards one of them is x8 that's no good. I'm I'm lo I'm like literally half the speed for that card interconnect. That sucks. So how do you get around this? Again it costs money. But I'm going to show you how it's done. So cuz I was really confused. I was like how how do you get around this? How come nobody makes just like a giant motherboard that has, you know, the PCIe slots spaced at what the spacing is that you would want and all that? It just doesn't really exist. But I asked around like, how are other people doing this? Because some people are running like 16 RTX Pro 6000s altogether. And it's like, well, how it's not possible, right? Um, so it's either MCIO or SlimSass. One of the things I found is this guy cain.com.
I'll put a link in the description. His name is like Christian Payne. some random dude in um no offense to him, uh random dude in Germany that started like I when I come to a website like this, I can immediately identify, oh, this is a shop page and I'm like this the this like this is some expensive stuff like a 2,000 euro switch or a thousand, you know, dollar or yeah, thousand euro um another version sort of of a switch. Um yeah, who is this guy? So, I did a little bit of research into this guy and like he he got his start doing like DIY risers on I forget what the forum was and then just kind of like went from there and he is very popular with uh people that have this exact problem. And I think that um you know he's just kind of doing it for fun and then now all of a sudden lots of people are trying to run local and I and I know like server grade hardware there are other ways to get this kind of equipment. Um, and then you could also go on Alibaba and find pretty much all of this and likely for much cheaper, but I always tr I try to use somebody who has a reputation to lose. Um, so I I you could probably buy it cheaper on Alibaba, but if that person um or that company just doesn't do a good job or even on like Amazon.com, this this is why I stopped using Amazon is you just have all these like random named companies and then if they start do, you know, selling a they sell usually they sell a good product and then eventually they start selling a crappy product and then once the reviews tank they just start a new a whole new company. So I like people who have a reputation to lose. I think Cain at this point has a a really great reputation.
So he has something to lose. I have not received my equipment yet. I have not used it. So I will let you guys know um when that happens, but I do want to give you a high level overview. So there's multiple ways that you can do this. One is if you just want the riser behavior and you don't actually need to get to like so like if you just want to use your slots and have risers but they work I believe the correct or the best way to do it now as far as I've heard this is still this can still be finicky but you could make use of slim sass. So these are all you've got slim sass and then you have like mcio and the way I believe to get like gen five speeds you you can't use slim sass anymore. you need to use MCIO, but I believe also MCIO is more reliable. If you're knowledgeable in this space, please feel free to comment below. I'm not an expert. I'm just telling you guys what I've learned.
So, um, the cheapest way to do it and just get risers if you're content with PCIe4 is Slim SAS. And then what you need is something like, um, uh, I don't think we need the rettimer.
All I want, let me click, let's just do I'm going to click Hey, I'll click these three things. Um, so first you would have like a passive host adapter and I forget the red driver. I was asking this guy like I you can email this guy. He's very good to respond by the way. Um, I forget what a red there's like red driver and rettimer and there are different things but you would likely just need a passive host adapter for slim sass if you're just trying to rise right and you're not trying to go to like a switch I want to say. So, um, you would have this. These are slim sass.
You have two slim sass cables that then would go, let's see if we can find one.
One of these. Let's see what's the difference between these two. X. Oh, this is an X4. And then this would be X8. Okay. So, then you could go to like one of these. So, again, two cables into one of these and then that your GPU would plug into that. So, that's a really that's a much cheaper. So, it's like 40 then you need two slim size cables and then 30. So, like €7 and then whatever the slim size cables are. And if you've bought risers, you know that yes, that's a little more expensive than a riser, but that should actually work as opposed to the risers not working.
So, one is you just piss your money away. One is it works. But then there is the if you want PCIe 5 speeds, you need to go to MCIO.
So, we come here. Uh, and then I'm going to click click click. [snorts] Uh, and then in this case, uh, I believe. So, you could pro I'm not even sure. Uh, this is sketchy to me.
Uh, where is the passive MCI5? I'm going to skip because I I already know I want to do timers. So, I don't want to talk out of my ass of you can actually get away with passive, but it seems to suggest that yes, you could plug this in. This would be a passive way. Okay.
And then once again, you would have MCIO and then you would go to your adapter and then plug in your card. But I think PCIe 5 is much more finicky and I probably wouldn't bother with it. So instead, the the the way the best way to go is one I'm going to speak in terms of four cards at a time.
So one or two rettimers depending on what you want. Um you can start with one and then go to two eventually. So, rettimer that go that goes into your actual motherboard PCIe slot. Then you have these two MCIO um cables that will plug into here and here. Those will feed to was it this? Is that what I'm thinking?
Probably. Yeah, that will feed to those cables will then feed to this. So, it's a retire. No, I'm sorry. I already screwed up. Well, it could do it could do it this way if you just want one thing there, right? if you just want to go to one card. But if you actually want to go to multiple cards, let me see if I can find my way back to sanity here.
Let's see if this is the right thing. I tried to save them. No, I' i've lost everything. It's all gone. Um, okay. You would do Let me just order I'm going to order this and then we'll we'll talk about it. Okay, I think I'm prepared. So you have t one or two timers per four cards and then this using MCIO cables feeds to your switch and then your switch has MCIO out so more MCIO cables to um your Gen 5 adapters. So you could plug four cards per switch and then theoretically you could have four cards that only take up a single PCIe slot or if you have two if that's if you just have one rettimer but allegedly you can get a little more speed um and throughput with two ret timers that's like the best way to to do it but you could do it with just one.
So that is how you do that. And essentially like the PCI switch, you can kind of think of it, you can also convert like if you have a PCI 3 or 4 motherboard, you can send you you can communicate with your your CPU.
You can send your communication from a switch to um the CPU over PCIe3. So you can still do that. But if you want to get gen 5 speed on an older motherboard, you can also use a gen 5 switch and then it's you can almost think of it like u like envy link or something like so your interconnect between those GPUs would be gen 5 even though your motherboard is gen 3. That is my belief. Um I could be mistaken. I don't think I am. But the only time you would really want you would really want your motherboard to be gen 5 is if you actually want maybe these guys these cards these four cards to communicate with another four cards.
then you would really really want that that like like that throughput between those two little clusters of cards. You want that to be really fast. So hopefully I've done it a decent amount of justice, but I just want to bring um that point home. Uh next, I have also upgraded my um my UPS, which is also my UPS serves two purposes. Brown out for my uh NAS. I don't want to corrupt my NAS, but then also you can use a UPS as a little more uh robust. It's not necessarily it's not a surge protector per se, but it separates your devices from the grid, especially if you have a nice UPS. So, my old UPS did not do this. It did not have this like double online conversion stuff. This UPS is a more appropriate UPS that literally does like isolate your um your the hardware that's plugged into these UPS is like truly isolated from the grid. Now, if we still took a direct lightning strike, you could still theoretically have an arc that just bridges pretty much any gap and you can still fry everything.
But in this case, this is a very nice UPS that uh is like 4 and a half to 5 kilowatts. So that that can power your AI server. You probably kind of need like one per one or two of those AI servers. And these things get expensive fast. So, the last thing I want to show you guys, if anybody's following in my footsteps and you live in like I live in I believe I'm actually in the worst place in all of the United States for lightning strikes, like we get constant lightning. I have had um I've blown through surge protectors. I've actually killed a couple of my switches, my like networking switches cuz I I didn't even think about it, but yeah, that's that's like a it's it's copper. Um, and then we also I've like I've got a solar array here and like a 200 foot, which is like I don't know 65 70 uh meter run of PV wire. It is in conduit and all that, but it's still, you know, effectively a giant antenna. And then I've got underground buried uh Ethernet cable, which is like shielded and all that stuff, but it's still it still apparently attracts. I mean, it's obvious now. And don't even comment. I don't even want to hear it. I should have buried fiber. If I ever next time I do the trenching, I will bury fiber. But for now, I just have like these like SPDs on both both ends of like my PV run, my Ethernet run, and ever since I did the SPDs, um it's it's been quite good. Um but anyway, I get a lot of lightning. So, uh and really even if you don't get a lot of lightning, at some point if you are doing builds like this, um it's very expensive. And these UPS's are very also very expensive. Like if you ever looked into UPS's um it is usually pro prohibitive enough that you're like I'll just take my chances.
But I also again I'm not sponsored by Cain or this other website I'm going to show you. But I found this website Greenlight UPS. I'll put them in the description too. They sell uh pre-owned and refurbished UPS systems. And in general I would say like server quality grade UPS like these uninterruptible power supplies.
There is nothing wrong. The only thing that really does kind of degrade over time, obviously all electronics, like I wouldn't want like a 30-year-old UPS, but a lot of these are like you could find them such that they couldn't like it's a model that couldn't possibly be that old, but then also it's the battery that decays over time. And all at least these guys, every UPS that you buy is has a replacement battery. So that's pretty cool. And it's like half off, like half the price. So it starts to make a lot of sense. Um, and they shipped like super fast. Like I couldn't I was like waiting for it to ship and then I got a call from the freight people that they're like we're going to deliver this tomorrow. I was like, "Oh, okay." [laughter] And so anyway, um yeah, cool. If you're looking for UPS, check out Greenlight UPS. I'll put them in the description.
Again, not sponsored, just really happy with um the service there. Uh and again, these people have unbelievable contact.
Like they're they're uh very fast to respond to you. So, moving on. So, uh, besides DSpark, I want to talk about some of these recent releases as well.
So, we had Kim K3 came out and I'm just like, oops, I'm just searching Kim K3 on on X. So, I'm going to scroll. So, Kim K3, if GLM52 was a worthy replacement of Opus 48, Kim K3 is a worthy replacement of Fable 5 and GPT 56 Soul. Like, it competes with those models. And um and again, it is theoretically an openweight model. They are set to release. Let's see if they say it somewhere. I'm not sure where they're going to say it, but uh in fact, let's go to their profile.
I'm It's got to be like pinned or something. Yeah, here we go. So, it's 2.8 trillion parameters. That's a big model. Like, initially, I thought, oh, I'll be I could at least run it on GPU and RAM, right? Because I have a terabyte of RAM. It's still not enough.
a ter I have a terabyte of RAM and like almost 300 gigabytes of VRAM. So I'm thinking like oh I could run Kim K3. No because even at 4bit it's going to be 1.4 terabytes.
[laughter] So it's like it's not possible man. This model it's just too big. Um so that kind of sucks. I don't think I will ever really run Kimi uh unfortunately. Like I could I can add more cards here and if I add my Switch I've got a lot of old GPUs so I could I've got my RTX Pro uh RTX 8000. So each of those is 48 gigs. So together and they have an envy link to each other. So uh those two cards can effectively be like one giant card. Uh and that's another 96 gigs. So I could have one more RTX Pro 6000 and then I could start throwing in the the the RTX 8000s and then I've got three 3090s, a 4090 and then a bunch of other cards. So theoretically I could ek my way to run Kimi, but Kim is going to run at like if even if I do that, it would run at probably like three or four tokens per second. I I I'd be shocked if I could get like 10. And I hear this model likes to use a lot of tokens. So I don't think I'm going to really run Kim K3, but it is cool that they release it.
Theoretically, it will be open weights.
We'll see if they actually release on July 27th, but I'm pretty confident that they will. Uh, and then moving along, we had Quinn today released another this one, another big gigantic freaking model, 2.4 trillion parameters. I have not um I have not seen their uh where are the benchmarks for this model? I don't even know. Show me the benchmarks, y'all. But so 2.4 trillion at um 4bit means that would be like 1.2 2 terb and that would leave us 200 gigs for context. So this is a model we really could um I mean even without using a switch that's a model we could run on this computer. So maybe before I get my switches I could try to run Quinn 38 if they really do release it. And this is another one I believe they said they will release the weights but they're not here yet. Yeah. So you can you can test it now.
I swear they said something about open weights.
I don't know. Anyway, um I might be mistaken. Yeah, it is going to open way soon. Okay, so um so there's that and then uh let me see if they released any benchmarks. Is this a benchmark agent world? Okay, these are not benchmarks. I want to see you know what I want to see. Here we go. No, this is 37. So dang, they didn't even release the benchmarks yet, but other people will. I'm sure. I'm sure it's a very good model. That's all I'm going to say.
Uh let's see if it is.
Let's see if it's written anywhere.
Yeah, I don't I'm not seeing their um I'm surprised they didn't like release benchmarks really. It's got to be somewhere. I'm I'm probably just being super blind and missing it. Ancient world.
Yeah, I don't know. Anyway, moving along. Um, also, uh, yeah, this is kind of crazy. Like this is like the last like, you know, few months, last quarter basically, not quarter, a little more than a quarter, um, of releases. You had Opus 47. I can't believe like Opus 47 was like not even that long ago. Kim K26, GBD55, Deepseek, and you can kind of see that you always kind of go sort of in order. Kim seems to have come early. We are We are due for a Deepseek model basically. Uh, for some reason I I thought I had Inkling. Let me just search Inkling real quick. Inkling.
So, thinking machines. So, this is this is Mera Marotti, ex OpenAI uh also OpenAI CEO for like a day [laughter] um back when that was an interesting thing that occurred. Uh so this is her company and I don't think anybody saw this coming that they were going to release openweight models and so they released a one Oh, I did have it open maybe. Yeah, they released a one trillion parameter model. Um, and this is already open weights. I'm trying to find like some benchmarks on this. I don't I'm not seeing some benchmarks yet. Uh, I'm just going to keep scrolling to the small one. Oh, here's some benchmarks for sure. Um, and if I'm honest, I'm going to say this model is not very good at it's not worth the size. Like, it's a it's a trillion parameter model. It's it's so big. And then you can start to compare this to models that are much smaller and it just isn't good enough to run I don't think like like terminal bench 21 638 is not good for a trillion parameters like that's crazy. So um that model is not very good in my opinion. No offense to these people, uh, and I appreciate the open weights release, but Inkling small, however, which is kind of funny because this one is a 276 billion parameter, so it's closer to like DC V4 flash. This model actually is like super similar in performance at a fourth of the size, a little le a little less than a fourth, but or a little more than a fourth, but is like pretty much the same scores all the way down the down the chart, right?
And this is kind of true of like DeepSc V4 Flash and DeepSc V4 Pro where they are so similar in performance, but Deepsec V4 Flash is so much faster that I do not know why you would ever use it.
And if you're over API, it's so much cheaper to use Deep Seek V4 Flash. You might not get all the speed because everyone pretty much always targets that 50 to 80 tokens a second, but it's so much cheaper. Like why would you ever use the bigger model? I do not know because like this this is like this might as well be the same performance across the board in my opinion.
[laughter] Um so small is not out yet there. It's in preview and allegedly this is on the way. I'm actually somewhat excited for Inkling small. Although again it um it looks really good compared to the full size Inkling. But even at this size it's I I I might still find Deep Seek V4 Flash to be a better model. But I'm excited to try this one. Um I think the thing with Deepseek V4 flash is it is a native precision mostly 4-bit model and that man that just changes everything.
Uh because that is like the benchmarks that you see for that model that is what you will get when you run it local and when you run it local you are running you're not running some quant you're not running some compression you are running that model in native precision and that's crazy. So uh so that was Inkling.
I'm trying to think if there's any other releases. Oh there was Grock 45 that came out. Let's see here's Grock. Uh here it is here on this uh list. And so here we can also see uh here's Keen K3 like right there. It might as well just be chilling with the Fable 5 and GPT56 Soul. Uh we still have GLM52 down here.
So even GLM52 is very close to these guys. But yes, for the Fable class model, Kim K3, but I would personally argue I don't see why I don't think it's worth the money for any of these models.
Like so 56 soul, Kim K3, Fable 5.
I don't think these models are really ever worth running. I think everybody thinks that they're working on the hardest, most challenging problems and they need the best AI in the world. Uh I reject that. I don't think that's true.
I just don't think it's true. Um so, uh I think I think what's happened to us as we've used these agents over time is they the models have gotten bigger and bigger and bigger. like we we expect Opus 48 is probably 3 to 5 trillion somewhere in that range. We expect Fable is likely 5 to 8 trillion parameters.
These models are just huge. So we come back to that time diverse token stuff and we come back to um the tokens per second max gen that you could make and all these these original problems. And the thing is it's so expensive to use and run these models. And I'm going to argue even if you're not running local, the price that you pay to use DeepS V4 Flash is so cheap and the intelligence, it's a dumber model. Like, yeah, sure, I can see that like Deep Seek V4 Flash 56 instead of 76 or whatever or 77. Um, and obviously benchmarks are, you know, they're just a silly number. But the other thing is benchmarks are pretty much all these benchmarks are a score of human out of the loop and I think what has happened is as these models have gotten bigger and bigger and slower and slower like when you're using like now when I'm like watching literally anybody use Opus 48 or Fable like if I'm if I'm sitting there and like watching someone try to like code with these models it is so painful because these models are so low.
And I think what is happening is we have been we're being conditioned to use these gigantic models that cost an arm and a leg, but they're so slow that you get a little distracted when you're using them. So either you're running 40 sub agents, which you're not, you're just not. Um, and if you are, whatever those sub agents are doing, they are not doing Fable 5 uh level work. Okay. Um, so either you're running multiple, you're trying to multitask, but you're probably doing that bad. Like humans, despite how badly we want to be good at multitasking, um, it is very hard. We are not good at it. Uh, so maybe you're multitasking, but I think I think for most people, at least when I like I'm just going off of what I have personally experienced myself and what I see around is people tend to actually just have like one thread of communication for real. Sometimes it fans out into sub agents for a temporary time, but most of the time it's like one main thread, one main task that you're trying to achieve with your model. And because they're so slow, people get distracted and they go do something else. Like they can't actually sit there and watch the model.
So they actually need a model that is human out of the loop. Like they don't want to actually be there. Not because they're stupid or they don't want to be an engineer anymore necessarily.
Not for all of them anyways. Um, it's because it is so boring and in the model is so slow and you just can't sit there and take part and keep your attention. So you need a model that is like fable that you can just do SLG goal and you go to bed and you hope you wake up in the morning and it's done for you or that you can just say do it and you just sit there. But it's like I like when when I really watch people do it, it's like you're just it's like that classic meme of like the old people playing the slot machines. That is what it looks like to me where it's it's like you're just kind of hitting enter and then you're waiting like 5 minutes and then you're typing a new thing. You're hitting enter and you're waiting another 5 10 minutes.
It just is so terrible. It's a terrible experience. It's not fun. And at least for me, I have not used Fable or GBD56 at, you know, at large yet. Um I did use Fable for a moment when it was like first released before it got pulled and I I it just didn't strike me as that great of a model and yet it's so much bigger. So, it's a little slower. It costs so much more. It just destroys your quota, right? Um, and I don't see the point. And I think like when I talk about this with people, they get a little offended because they're like, "Oh, well, yeah, but it's so good." What are you What are you doing? What are you doing? Because literally like like four months ago, five months ago, everybody was getting by with like Opus 45 and everybody was using Opus 45 because they thought that's what they needed to do their job. And then now all of a sudden people think they need Fable and GPT56 Soul to do their job.
What happened? What are you doing? I don't think so. I think everybody's doing SAS companies. So shut up that you think you're working on a problem. is so hard. You need Fable or GBD56 Soul. You don't. Um I think, like I said, I think the reason why people actually feel like they do need that is because the model is so slow. And when you're using something like um like DeepSseek V4 flash, especially if you can run that local and get that extreme speed out of it, I actually think that model becomes more useful, more productive and a better model than at least my experience on GPT or I'm sorry, GPD48 uh G Opus 48. That was what I you that was like the latest biggest model that I really got a lot of use out of. And I think I think people are making the wrong calculation at this point. Like I don't think you need Fable uh to do most of your day-to-day. Now maybe there are some things like like sometimes people are writing you know custom kernels to do really novel and out there niche things and having a really smart model can help with that kind of stuff especially if you really are like completely out of your zone. um and that is really truly what you're doing or you really are not an engineer, not a software developer and you really do again want to be way outside your your comfort zone. Um these these giant models can do that for you. But I like my experience on Deep Seek V4 Flash is it also does that and it's just so much faster. It iterates so quick and it it solves its own problems like so fast. As long as you're willing to take part like human in the loop, I think it's just as smart as any other model I've ever used, including Fable 5. And I know that's going to piss some people off, but the problem is it's it's like human in the loop versus not in the loop. Do you want to take part? Like I think I think what I'll close on is I don't actually want my AI model to think for me. I want my AI model to be a tool that I use, an extension of me. And I think as time has gone on, we have slowly like relegated the actual thinking, the actual planning, the engineering, the problem solving skills, the inferring what might be a problem down the line, like all these things. I think we're we are offloading that to the AI more and more. And that's why we feel like we need these bigger models.
But and it's true like I said it's so painful to use these models that are like because they're so slow and because like I was explaining even like with the together API sometimes it's 200 tokens per second sometimes five and you never know like when you make that request you don't know how fast it's going to be. So it's kind of it's just it's it's uncomfortable to sit there and just wait like you're just staring at text scrolling on the screen. That sucks.
It's not fun. It's not engaging in any way. Um so I get it. I get why people don't like that, but I think that's just another reason why I I think local is very powerful. And I don't actually think for many people who are software engineers, this is how you make your living. I don't think owning two RTX Pro 6000s is crazy. Um I don't I just don't I know it's expensive, but it's I don't know. I I think that it's a a worthy pursuit if this is how you make your living. So, um anyway, yeah, I think that's it. And yeah, the main the main thing is I think I think it's a mindset of outsourcing your thinking, letting a model thing for you versus being an extension of you and a tool that you use. And if you're looking for AI, like if you feel what I've just described, um, if you're looking for something that doesn't feel so hollow and and like as you use it, like you're not actually doing anything, I I would propose that at least start trying to use Deepseek V4 Flash through an API and just see is this model really that bad? Because it's really cheap. You can use that over API and you'll have a hard time hitting $200 a month um, expense. Um, try just see just see what you think. Um, and then also the other thing I want to talk momentarily about, I forgot is yeah, Grock 45 came out and it's not, you know, Fable class necessarily, but it's pretty close. It's a strong model. But the only thing I can remember about 45 coming out is they also released Grock Build, which is like their CLI. Now, if you want a CLI that does not track anything that you do, you should definitely check out Minion on GitHub. Let's see, is this my Minion?
Yeah, I'll put a link to Minion in the description as well. Um, but anyway, um, Grock 45 came out and, uh, they released Grock Build and then what Grock Build was doing, let me see if I can find, I think I have a Reddit page. Allegedly, Grock 45 or like or really Grock build CLI was uploading your entire repo and git history and apparently. M to XAI's cloud, which is kind of crazy. That is that is legitimately crazy. Um, and you couldn't stop it apparently. And I do think that's very egregious and I agree.
And allegedly they have fixed this.
But even if they've fixed it, I want again as I make my case for running local AI, I want to draw your attention to the fact that this sounds really bad and for them to just do it all in like one fell swoops does sound bad. But when you're using any API when you um like I have tried really hard to like not let M leak and or like for not like things inside of M to be sent over but if you actually pay a lot of attention this is that that could be kind of challenging but not just M secrets I'm talking like your entire like company your all your IP all the way that you think everything when you're interacting with a when you ask for help on like a um a piece of software, what do you think happens? It sends the entire file as context or at least a large part of that file. But generally, like if you really want to know what a file is doing or what a piece of software is doing, you have to look at all the files, look at what's in there, how they interact with each other, you have to reason over that. What is that? It's sending all of that text data as context to the API.
So even though this sounds really bad and it does that is what is happening when you are using claude or chatgbt like all these companies have all of this and more it's not just your your you know your env secrets it's it's your entire company it's the way that you think it's all the code that you've ever written or that you've had written for you it's it's literally everything and they have all of it and so again it's like you want local because privacy and not just it's not like you want privacy because you're doing anything wrong or you want to ask sketchy questions, it's because no, I actually just don't want to upload everything I do to these companies and for them to train on it and steal my company, you know, stuff like that. And again, I think everybody sees this and and they think, "Oh my god, like that's terrible." But like it's been happening this whole time.
[laughter] It's just this was like the most egregious and the most in your face. But it's been going on for a while. So, u anyway, uh that's all for now. I will put a link in the description for this.
I do have once again the full um the full breakdown of like all the scores and the performance and all that. Um and then all the throughput stuff. Long story short, Deep Sync V4 Flash DS-Spark amazing model. I am looking forward to seeing what they put out next. Um we're any day now they're going to put out a model. So I'm curious what they do next.
Um it's going to be hard to dethrone DeepC V4 Flash. There's a couple other models coming out that I'm aware of. Um so I'm excited for those. I know of at least one that kind of competes with Deep Seek V4 Flash, but it's like half the size. So, um maybe more on that when they finally release that model. And uh yeah, exciting times. Like so much is happening. And the the craziest part truly is still there's no moat. Kim K3 came out.
You can't say Kim K3 was a distillation because it has to have been pretty much done by the time 56 Soul came out and then Fable 5 was like inaccessible for a while. you just can't argue it anymore, you know, like so you can't just say, "Oh, they just distilled." And also, like, let's be clear, every model on this chart distills either themselves, other model, like you'd be stupid to not distill. Okay? So, like I I I think we need to stop with this nonsense about distill is is somehow bad or evil. And then also all everyone on this entire chart, there ain't a single a single entity other than I don't know. I have to double check Neotron 3 Ultra how their their data because they do release their data. everyone else on here, they are training on data that they potentially don't have the rights to.
Like they're pretty pretty much everybody on this list allegedly uh I think the argument can be made they're committing uh copyright infringement.
Okay, so shut up about distillation. I don't want to hear it from at least the whenever these companies are crying about distillation, shut up. [laughter] Just shut up. It's not your data. You do not own those outputs. I understand we we're paying you for your your training time, your inference serving. That's what we pay for. Stop trying to claim ownership of that data. Uh I don't I don't like where that heads. Um so anyway, okay, cool. Uh long story short, risers suck. We hate lightning. Um D V4 Flash with DSpark is king. And I can't wait to see what all has come out the next time I apparently record a video because man, things are heating up fast.
Things are going quick. Um, all right.
See you guys in another video.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

WOW! Judge TURNS THE TABLES on Trump in His OWN $10B LAWSUIT!!!
MeidasTouch
197K views•2026-07-23

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Steam and Xbox Just Dropped The Hammer On PlayStation
OhNoItsAlexx
9K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23