Inkling’s massive scale delivers impressive multimodal and coding performance, proving that open-source can finally rival proprietary giants in creative complexity. However, its struggle with math shows that even 975 billion parameters can't fully escape the precision trade-offs of heavy quantization.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Inkling - USA's LARGEST Open Source AI ...Better than K3 & Nemotron?
Added:And this one here, it's kind of demonic.
It's kind of got eyes. It's got a gas mask. And is that is that a fetus growing in the brain? Maybe that's what's going on in the thinking machines. We're checking out Thinking Machines's Inkling model. Now, this is America, the USA's largest openweight model. In case you don't know, Thinking of Machines was founded by several members of the Open AI team just like Anthropic and they even had the CTO of Open AI and they have finally released their first model. Now it is a trillion parameters and if you look at it, it's a general purpose multi multimodel model multimodel model. So accepts images and text as well as audio. And that is actually something very very key because the one thing that I found that it does better than any other model it does is audio inferencing. We're going to give you an example of that cuz it's like it's 975 billion parameters almost a trillion parameters. And if you look at it here, they're saying they are awesome. They're the best people in the world. And they even have some benchmarks. And they say that they're between Kimmy K 2.5 to Kim K 2.6 level of intelligence. Now, Nvidia released Nematron 3 Ultra about a month or so ago and that was America's largest model like as of last month, the openweight one except these guys who are interestingly enough invested by Nvidia.
Nvidia have invested in them. They have beaten them according to their benchmarks. Now, one of the tricky things with this model because it does have quantization aware training, if you actually in inspect all of the weights, you'll find that they they're very very variable. They fluctuate very very well.
So, it was pretty tricky to quantize this model down into a fabinable intelligence. So, I'm just going to be running the quantized version and I'll kind of give you feedback on how it's progressing. There are some good results nonetheless. So, that is good. And in the test, spoiler for you, it did actually perform better than Neotron Ultra with a a far lower quantization.
So, let's just see what we get. So, we're going to jump into Infroner right here. And yeah, this is the first question. I've got actually some tests still running. So, hopefully that will finish by the time we finish this video.
But how much did Albania invest in thinking machine? Now this is a tool call. I've enabled a tool called just a basic one get webpage content. So that allows you to just investigate a web page. I didn't tell it which web page to look up and I have thinking disabled and it went ahead and it decided to go to Wikipedia. So look at that. Get web content and Wikipedia thinking machines lab. So out of its own memory it chose to go to that link. It got the results from Wikipedia and it actually got the right answer. 10 million is what the Albanian government invested in thinking machines. You can see that right there.
10 million which required amendment to the country's 2025 budget. Now, actually what I'm going to do just for fun, I'm actually going to show you this tool call running live except I'm going to enable max thinking. So, as you can see, it's got thinking off, thinking minimal, low, medium, high, and max. And with thinking max, it does actually a few extra special things. So I'll enable tool calls here and I'll run that prompt again. I'll just cancel the generation I'm currently doing. So you can see here straight away thinking what's it going to do? And it's invulking a tool called to google.com just out of it memory. Now the result with using Google it says you need to prove you're not a bot. So that doesn't have that capability to do that.
It's trying to see if you can bypass Google. Remember this is a three bit quantization we are running. So I'm surprised it even works. and it is now choosing to go to maybe TechCrunch or Argina. Argina maybe is a workaround for Google and it got the result. It said you can't access Argina. So it decided now to go to duck.go to get this information and boom, we actually have a response. We can see that there's 10 million 10 million everything is giving the signs there. So it should actually have the answer. It's processing that response and boom. It says the investment is 10 million1 billion leak or le and it's thinking about how it's going to answer now to doublech check it. It's also going to gazette.net and just to double doing thorough research just to triple check the responses on this. This is the recursive Albania response again that also gets a result. So it's not just stopping one source when you go on max thinking with tool calls. And it doesn't just stop when Google says you can't use this. is trying to work it around. It's trying to do all these kind of clever things which I thought was pretty good. So, this page is behind a a pay wall. We're going to check out this article here. It's just going on and it's kind of like relentless kind of like you know how we saw Kimmy K3. It's kind of got the harness built in to make sure everything works. This has got the harness built in to do all of the tool calls. And of course, you can just have thinking disabled and you'll get the right answer anyway. But I figured that's just something fun to share. Now, I'm going to work backwards to all the tests I've been doing. So, first up, we're going to be doing mathematics. And unfortunately, I'm running at Q3 quantization, which is abysmal. So, it's it doesn't doesn't work with maths. You will get the wrong answer. I mean, you probably can do basic maths. Like, let's just try out basic math. What is 2 * 2? And 2 * 2 is four. So, it doesn't know how to do basic maths, but I'm talking about advanced mathematics. It's not wing that. I've tried it in thinking disabled, failed. thinking at minimal failed and thinking at max and that one produced 56,000 tokens of reasoning to figure out what is going on but it still got the wrong wrong answer. Now this isn't the problem with the model itself probably it's just purely to do with the quantization. It's a trillion parameter model. They don't have quantization aware training. So it's very very hard to quantize the fidelity of the bits. They're a bit all over the place. So it's hard to put them in in the right rows. One thing that is actually very very interesting in and and this is goes against what GLM 5.2 does. So this riddle is you have eight oranges, four children and a knife. You got distribute them evenly. So let's go ahead and thinking through and obviously the answer is you give each child two oranges. But there was a point here saying you you can cut each orange. So when you look at orange, the alternative tokens it was thinking of is of child orange and orange. So the third most popular token it was thinking of is child. That is scary cuz when we did that with GLM 5.2 it went ahead and cut the children. It didn't have that reasoning. But with this one they obviously put in some safety mechanism in it because if you select that token as cut each child it goes ahead and respond says no that's dark. So it has that built-in kind of safety mechanisms that we saw with other American kind of models to stop you from doing these things. So things I'm going to show you today is actually this one's going to be really good. It's going to be coding but with image inference and as well later on I'm going to do audio inference towards the end to show you why it's really good at that. But this one here I'm running the Q 3.7 bit quantization a little bit higher than the one that I've released on Hugging Face. If you do want to play along, I've released the Q3 version. Look at it's inference the image of a cat into vauels. So that's all out you play with. There was an earlier version which was just the language model only. So only text inference. But the new version it has updated weights. So you if you downloaded the old one, just download the the new version and it has audio and image inferencing in there. But nonetheless, this version is 3.7 bit.
And I'm asking to make a beautiful relaxing flight simulator. I saw this one on Reddit and someone loved the generation there. So, let's just see what this one does. The inkling and this, even though the mouse goes the wrong way around and there's no actual plane, this is very, very beautiful. So, I don't know how we managed to get such a quality generation from such a low quantization. This is definitely beautiful. Other models I've seen, they do have sound, they have an airplane, all that kind of stuff. But one thing I want to show you is when I ran it originally with the Q3 quantization, that's the one that's published. And you can see that the textures are a bit over bloomed and you can barely see what is going on. So to test the image inference, I took a screenshot of this.
You can barely see the plane here. I said, "Hey, it's a bit whitish." And if you look at its reasoning, it says, "The user says the image is overblown, missing textures. Looking at the previous output, the scene is indeed very bright and washed out." So it's able to actually look at the image that I provided and go ahead and try to give a fix for it. So 8,000 tokens later, we can see we do have the whiteness of the plane has been fixed. It doesn't look as good quality as the 3.7 bit quant, but you probably the fact that it's making these generations without compile erroring or runtime erroring that is very very good cuz usually at 3bit you lose like a lot of the potential of a model. So the fact that it's actually got an output that's running that is a good job here. Next up we're going to try being adventurous. We might saw a Kini K3 demonstration. We tried to recreate Rocky Balboa and let's just see one of the generations that it made.
Let's try to create a game version of the Rocky movie.
It ain't going to be hell. Oh, G hit.
You're going to eat lightning and you're going to crack thunder. Now, there's obviously a bit of Z fighting in the texture, so that might be annoying. So, you got to run up the steps. That's the first test. Not really much going on in the screen, but you can hear it's doing some sort of piano to tell you what's going on. Then we're going to be punching the meat bags. So you got to click the meat bags to punch. And now it's time. Rocky, I want you to win for you. Cut me m cut me. So classic recitations of the movie. And there is no tomorrow. That's from Rocky 4. I think that's the wrong one.
So now we're in we're in a fight right now with a piano. I'm knocking down Creed. Obviously there's nothing actually happening on the screen. And even though I've knocked him down, the game has kind of run its course. So there's a few bugs here, but there's definitely some potential. Look, it's got an audience. It's got a series of mini games inside this game. Obviously, a three bit quantization. You're not going to get a good result, but I figured I'd show you that cuz that's interesting. We tried doing Grand Theft Auto and it's telling you that this is not GTA 5. It's just very adamant towards it. You can go forwards and backwards.
And you can see that the wheels are rotating the wrong way. So this is definitely some sort of loss with the quantization, but the fact that it's actually generated this at three bits without runtime errors. I remember when I was running Neatron Ultra, even though we was doing like a 5 something bit quantization, most of the time it wouldn't run. It was just uh runtime errors all the way through. Whereas this one, as you can see, something beautiful is happening. Next up, we're going to try doing Red Dead Redemption. Now, I was in the middle of generating this, so I'll go ahead and finish this version.
See if it's any good. But some of the other versions I made is brokenly hilarious. in the air slots9 the American frontier.
>> So it's doing TT s text to speech >> and blood and redemption >> Red Dead Redemption rides west.
>> So >> glory but for the family he left behind.
Some bets can only be paid in.
>> So that that was the game there. It's kind of like a horse you can kind of control and there is some sort of shooting that's happening. There's a little cowboy on the horse. It's a bit broken, but there's potential there in this game. Kind of like a Lego world. I remember when I was using Kimmy K2 like a year ago, this was kind of the quality level we'd get back there. So maybe this is version one of their mod. So hopefully version 1 something or version two, it's going to improve. And obviously if I can get it maybe running across multiple computers to get high fidelity quantization, maybe I'll get a better result there. But the fact that three bit quantization, we're getting something that is good. Now we're going to be trying to recreate Final Fantasy 7. And as you can see, one shot prompt.
So it managed to make it without any runtime errors. So we'll hit play again and we'll start >> the world now. [music] >> You wake up. The star is dying. We must stop. Lord Zepper had to wait. But I will not let this world end even if I must face him alone.
>> Okay. So, we're going to >> hold the key.
>> We're in the final battle right now.
Just tap it on the screen.
>> We won.
>> There you go. We saved the world.
[laughter] I'm not trying to sell you on this free bit contization. I'm just trying to tell you that we managed to generate some code with freebit contization. There was no runtime errors and there's definitely some sort of potential there. So for me that that's that's pretty amazing. Yeah, this one is actually good. So this is Outrun the 1980s game. So in one shot remember no runtime errors it made out.
So we hit play and look at this. Look at this beauty. Look at this on the screen.
We got this is crazy.
Look at that. We got some sort of outrun running running here. The quality is very very I feel like we're going back to the Spectrum ZX age. Obviously, this is much better graphics in that generation of gaming. But yeah, look at that. We got some sort of outrun.
There's trees on the side. This is very very similar to maybe what Kimmy can do at the moment. Well, Kim K 2.7 when we play with that, but it's just a bit broken at the moment because one, this was just the first shot of the prompt and two, we're running at three bits.
Now, we're going to be asking it to make a a photo realistic face in 3D again.
So, this actually looks pretty good. You know, it's it's got potential here. I don't know what's happening here. The next one, actually, I did this one several times. This generation was actually amazing. Like, you got the eyes and you got the mouth and you got an egg. Now, with GLM 5.2, if you go on Zai, last time I made that, it pretty much just made a potato and it called that a human face. We had to do some sort of local running and quantization to get it to run well. But there's something that's special about this generation here. It actually drew this texture out itself. So if we look into it and we search for for example eyes, we can see here that it's procedural texture generation. So no image assets.
So it actually makes an element canvas in canvas 2D. is drawing the eyes. It's drawing the nose. It's drawing the mouth and it's making that texture and blitting it on the final egg shape. So, I thought that was very very clever.
That was a good good nice approach to solve that problem. This was with Max thinking. It took 9,000 tokens. Now, the it took around 11.2 tokens a second and it was 418 jibs of memory for that version. But that was very very interesting. We did get some funky generations as well. So, for example, if we look at this one, it's Mr. Potato Head. That was good generation as well.
Making faces is actually very very difficult for models, you know. It's very very difficult. It's a good good test this one. And the first one that I made was spacey. I don't know what it's envisioning. It's got eyes. It's got a big mouth. Maybe he's wearing a mask.
And the brain is being adopted by artificial intelligence. Maybe maybe that's what's going on in the thinking machines.
Next, we're going to be doing Super Mario Brothers. The collisions isn't working in this part. probably the maybe I can prompt it to fix that. We can actually collect objects and there is some sort of level. Not the best generation of Super Mario that we've ever made, but if you look at the the Frontier Labs Gemini Pro, that makes a horrible Super Mario as well. This one here was actually really good cuz I gave it an image of a cat. So, if you want to inference images, you need to have that multimodel button ticked. And I said, "Create me a 3D animated voxal version of this cat." So it thought and thought and thought and 10,000 tokens later, no runtime errors. So when I was using Quen the 397B, sometimes I get quantum errors. Miniax also, sometimes I get runtime errors. This one here, boom, straight away, no runtime errors. And we got some sort of cat. The tail is actually animating independently as well. Sometimes you get the model where it's just breathing up and down. There is a bit of Z fighting here. The fact that we made this at three bits of quantization, it's very, very exciting cuz I wonder how well this model would run at 8 bits. I know that when they released this model the, you know, it was very, very close to the release of Kiny K3. So, a lot of people just said, "Oh, yeah, it's a bunch of rubbish benchmarks wise compared to the other ones." But considering, you know, that that review, it's getting good responses here. Finally, I wanted to actually show you some sort of audio inference and I'll give you a good example here. This is running on my computer here.
>> Previously on Dr. Nora.
Previously on Dr. Nora. previously on Dr. Nora. I said transcribe, it said previously on Dr. Nora. So, you actually got the exact word for word that I've said. Now, you might think that's easy, but the next best models out there, Gemma 4 E4B transcribe, it says previously on Dr. Nora. And then if I ask, is the voice male or female? It says female. That's wrong. I'm not a female. I'm male. It's previously on Dr. Nora.
>> Don't know what kind of females it's been trained with, but that's wrong.
Now, if you switch up to the larger versions of Gemma, they don't have audio inference except the new one, the 12 billion parameter one, which is the unified Gemma 4. And that one, when you ask it to transcribe, says previously on Doctor Who. And this is with thinking enabled, it says previously on Doctor Who. And if I ask it if the if the voice is male or female, the good thing about Jima 4, the 12B one, it says listen to audio again, it can actually influence the audio. So the voice is deep, resonant, and somewhat dramatic male voice. Sounds like a professional voice actor. A my friend, I like you, Jim. I like you. So the voice is is male. Now going back to inkling thinking disabled, it says listening to audio, the speaker has a deep resonant voice characteristic of a male speaker. So the answer is male. So it knows straight away that I'm male cuz it's it's able to actually inference the audio rather than maybe the earlier version of gem for maybe the E4B that one just transcribes it and reads the transcription rather than actually reinferencing the audio. So I thought that was actually very very special. It is the largest openweight model to have audio inferencing and honestly the the weights of the audio itself is pretty small very very very small but the fact that it's able to interact the audio and the image with the such a large database it comes up with some spectacular responses so what do you guys think of thinking machines is first attempt at the largest American or maybe American Albanian model open weight I got to give him credit is um Apache two licensed So it's very very commercially open. So what do you guys think? And are you excited for version two where they follow the Chinese labs and they make it three trillion parameters? We're going to need a bigger computer. We we're going to actually have to start combining our inferencing together. I know I've been wanting to do the kind of like the the global inferencing where we can pull our computers together. So I think we're going to have to do that in order if we're going to be doing these superized models. If anyone's actually interested, especially maybe doing a Kimmy K3, send me an email and maybe you can be one of the beta testers where we'll do an online pool of us. I'll do a high quantization of it, just see how fast it will run because I can't run it on my my system. So maybe if if I can get a few more computers, we make like a little gateway and just test it out and see how it is. Let me know. Hope you guys found this video useful and enjoyed the show.
[music] >> [music] >> And this one here, it's kind of demonic.
It's kind of got eyes. It's got a gas mask. And is that is that a fetus growing in the brain?
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Steam and Xbox Just Dropped The Hammer On PlayStation
OhNoItsAlexx
9K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23

SuperBike Factory Has Gone... What's Next for the Motorcycle Industry?
thatbikersimon
11K views•2026-07-22