Setting up a local LLM requires integrating an inference engine (like Llama.cpp), a framework (like Rig Core), and proper tool calling mechanisms; the key challenge is ensuring the server correctly intercepts and executes tool calls rather than returning raw JSON, which requires using the Ginger chat template and proper API configuration.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Setting up a Local LLM - LIVE [3]
Added:Okay, we are live.
But now is a good time for me to test my stream chat application.
I think I'll be monitoring the chat, but I'll also try to use my thingy. Let's see. test.
Oh my god, it works.
It works.
Very nice.
Hello, NK nappy.
Hello Omare. Hello Jason.
Is the thing on? Yes.
It's my pleasure to announce It is my pleasure to announce that my stream chat application is working. Look at this.
Hello Jack. Hello Mash big Hey.
Hello Tolga.
Hello.
Yes. Yes. Video on my Linux setup.
What's specifically about my Linux setup?
Someone did ask about um like showing my desk setup.
Hello Nathan.
Hello.
I could show my desk setup somehow.
Hello, Shristalin.
What can 8 gig of VRAM do?
Um, it's running on a Jetson though, which has a TPU.
I don't know exactly what it's capable of.
What AI do I use?
Um, Claude, Anti-gravity, sometimes Codeex.
Hello, Hidnut. Hello, Marco.
Do I have any sweet road map? I would just say go to roadmap.sh.
You can find a road map here for that.
The work has already been done software engineer. Let me see.
Where is it?
Backend. Front end.
Which one? Which one you looking for?
How does this crazy thing render on the terminal?
Well, it looks like that. Is that emojis [sighs] face blue smiling?
Let me see.
What the hell is Oh, it wasn't an emoji.
Uh, but that's what the emojis look like.
This is fantastic. I'm so glad my little program is working.
Hello, Henry.
Uniform text.
What is that?
Ununiform numeric sign 8.
It's like some pattern.
Ancient Mesopotamian uniform numeric sign.
It represents the number eight [snorts] in base 60 number system.
Whoa.
What?
That's crazy.
Base 60.
is that that's so strange.
[snorts] Yaml, they were ahead of their times. Yeah, apparently.
This is like some ancient mystical thing.
Am I renting the GPUs in AWS? No, I've got one locally.
All right. I I need to I need to look into this later.
It's going to pin that and look into that.
But in the meantime, the plan for today is uh I wanted to do something in particular.
Nazis put us on the moon. Never forget.
Um, so I want to set up my Jetson, which by the way, I have got [snorts] my Jetson running on my local network uh with its uh models here. It's only got one model at the moment, which is Quen 2.5 3 billion. [snorts] um uh parameters and I've actually got PI agent hooked up to it.
[snorts] I don't know if you guys probably can't read that, right? I need to make that bigger. So yeah, I can hook up Pi to talk to my local Jetson.
Um [snorts] Oh, wow. So, it said yes, I can fetch web pages. And then it said, I'm sorry, I'm not able to access the internet or browse web pages.
So, it's super consistent and very very effective.
[snorts] Uh, but essentially I wanted to get it hooked up to my Telegram and our chats coming through. Put it in production ASAP. Yeah.
So yeah, hooked it up to to Telegram so I can message it from my phone. And there's this thing called rig core which is like a rust library for building LLM powered applications.
Um so essentially I would set up rig core [snorts] probably it would talk to telegram and then rig core would talk to it would relay the messages to the model.
um and say do tasks. So maybe I could set up my own tools so that it can fetch web pages.
Bro, give away one to me. What? A Jetson?
I only have one. You cannot take my only Jetson.
Uh, that's crazy why it's uh I'm going to actually give this.
Let's see.
We're going to give this transcript to uh anti-gravity and see what it thinks of this result.
Let's see. I have a LLM running on my LAN at this address. address fried by agent and getting it to it said this.
Let's see what Gemini says about that.
Have I tried have I used Chinese models along with clawed uh with open code? No, I haven't tried that.
Gemini CLI is just dump.
[snorts] I don't have a problem with it.
Just tried it with open code. Go and I must say that it's all Is the text cut off? Awesome. Yeah, my text is not wrapping. Damn.
Will I stop YouTube if I get employed again? No, I probably won't stop.
Uh, you know, unless my employer is asking me to stop and also paying me lots of money, then maybe. But that would probably only happen if I joined a trading company that wanted to be very secretive about everything.
I'm trying Nvidia's Neotron models and they work pretty good with Hermes.
Neotron.
Never heard of it.
It's a family of open models with open weights. Oh, really?
That's kind of cool.
Oh, I can I I can grab this on on hugging face as well.
550 billion. 30 billion. These are all too big for my Jetson, I think.
Someone uncle, you're looking so handsome today. Thank you very much.
Someone called you uncle instead of pro.
What am I trying to build today? The goal is to build like um a sort of my own agent I suppose um with Telegram and this thing here rig core to get my local LLM to do tasks. But yeah, what keyword am I using? It's a new fee. Go to my website which is visarchus.dev.
dev/keyboard and there's links to to the I'm actually I'm using uh this top one here, the kick kick 75 with brown switches, low profile.
You tried it with open code pretty bad at the time. Made a mistake on simple commands. What model am I running on the Jetson right now? We're going with Quen 2.5 3 billion.
Um, yeah, it's pretty small. What do I think of Kimmy? I don't know. I haven't used it yet. Posted a short about it.
Yep, I did post a short about my keyboard. My own harness. I'm not sure if it's my own harness. I guess it's my own harness, isn't it? Yeah, I have member. I had no idea. Probably someone gifted you membership. It was probably B. Diddy that gifted you membership.
Have I tried Prisms yet? No. What's that?
Bonsai the model name.
So you hosted your site in GitHub. I also have a domain. Yes, I recently got a domain. Um, so now my my blog is set up with a uh custom domain, visarchus.dev.
It's nice and cheap. It was $10. Pretty cool.
So I myself tried setting up the local LM using LM Studio 4 billion parameter model on a 3050 GPU with 4 gig of RAM.
You struggled with the speed of response. Fair enough.
Large models can't fit on smartphones.
Data centers can't sustain them. Prism is building ultra dense intelligence.
Cool thing about Pi is you can ask it to configure itself and it will write extensions for itself. Oh, that's cool.
I need to fix up this uh this actually.
Can I rate your portfolio?
Um best thing to do there is probably to um is to send me an email and I'll get to that when I can.
lines if they edge.
Yeah. Let's see what it does there. Why don't I use aerospace? What do you mean?
What's aerospace?
What the hell did did anti-gravity do here?
It it went and implemented the the project window manager on Mac OS. I'm not on Mac OS at the moment.
This is Linux.
Still Gemini. Yes, I like I like Gemini.
What's the problem? [laughter] Please use Claude code. I use Claude all the time.
I honestly don't see the difference.
Did I get a chance to look at Laura? I briefly looked into uh uh Laura and that seems really interesting and I'm wondering how I can get that set up with my Jetson because that would in theory mean that I can do much um more demanding tasks with you know less less VRAM. So that would be good I think implemented terminal line wrapping calculated blah blah blah. Yeah, cool.
Let's see.
Gemini got degraded.
Okay, we have line wrapping. Very good.
Let's clear the chat.
what's going on but how to even encounter that problem like we can go with one billion parameter model as well for mostly day-to-day coding we need efficient tool calling let me pop my chat down here again I haven't tried Kimmy use something like vibe proxy which lets you connect all types of LLM to any type of harness editor, vibe proxy, Mac OS.
There's so many different vibe proxies here.
Use GLM.
I why GLM over Quen? Right now we're using Quen.
Yeah. So, okay, it just decided to go ahead and implement absolutely everything. That's crazy.
So it's added rig telixide Tokyo suras HTML to text.
What happened?
We've got a main in here.
Let's take a look.
Better to use Oama cloud sub.
Gemini makes my code base more [ __ ] Well, isn't that up to you if you review the code?
Anthropic bans you if you try to use your Claude subscription outside their products. I'm not subscribed with Claude at the moment, but I still use it.
Uh let's say on someone else's subscription.
So, what do we what what has Gemini actually done here? Let's just uh this text is massive, but it's probably readable.
[clears throat] We are doing what exactly?
How's the experience using local LM? I'm kind of getting set up, but so far it's a bit janky. If you take a look at um if I run pi and I say hello uh what model are you? It see it's sometimes a bit weird.
You know what I mean? Like what is this?
That's weird. How do you use claude code without subscribing? I'm using someone else's subscription.
in a commercial setting.
Put Ruhan on timeout. [laughter] He hasn't spammed yet, but spamming will be met with timeouts.
[clears throat] So, it has set up some logging. It's got a telside token. It's got an LLM API URL with a default to the local one. The model.
Um, then it has a client with no API key. Then it has this agent.
What's going on? Is there actual spam happening?
We're good. We're okay.
Um then it is saying it's giving a preamble. You are a helpful AI assistant called PI agent. You have the ability to fetch web pages using the fetch page tool. When the user asks about a website or requests you to look up a URL, you must this seems excessive and silly.
web fetch. So I can just give it whatever tools I want kind of. So here we have imple tool. Where is this trait?
Okay, rig core has this trait called tool and tool.
Let me see this modules tool. And here we have traits tool.
>> [clears throat] >> A simple LLM tool has a name, description, and parameters call. That's kind of cool.
I like this.
Is this an example? Yeah, this is an example.
Add X and Y together.
Blah blah blah. arguments self argu self args add argu okay it's a strct where did I find out about rig core gemini told me is the name son of jet inspired from son of Anton no that's funny though the name son of jet is literally just from Jetson so Nvidia Jetson I just inverted that son of jet But that is funny. Son of Anton rig is bad. It's very opinionated. I like opinionated.
You should create a tool for indexing an entire project for fast access to LLM.
Maybe.
Have I heard of conductor?
Allows you to use claw code and open code in parallel.
What scenario would I want to do that though?
Can't I just open like multiple terminal windows or something?
So we have web fetch where it is setting up a request client and doing the thing and it's giving back plain text. So it's actually converting the HTML to text.
That's interesting.
Convention over configuration.
What was that in response to in Atlassian? Did people use NeoVim? Uh, most people used VS Code, speedrunning token consumption.
Well, the goal is to not have to consume tokens from the big the big boys if I can avoid it.
Um, so that's that's the tool. That's one tool there, let's say.
Then we have Yeah, basically we're creating this client, giving it the tool. We've got this agent here. We've got a history map. I'm not sure what that's about.
It's a hashmap with chat ID which is what it's from here probably and then rig message.
Okay.
In support of opinionated libraries.
Uhhuh. What inference engine am I using?
Don't know to be honest. I don't even know what that means.
inference engine.
Oh. Uh I think I think the answer to that would be llama CPP.
I think that would be the answer there.
So we have what's this telide ripple um with histories and the agent. So for each message no chat ID user text if it's clear or reset.
Uhhuh.
Send message.
Have I tried MTP models? I'm not sure what MTP is.
Notify user that we are thinking of fetching. Huh.
Query the agent. Let's just I'm going to just kind of try to run this and see what happens. We have one expect here.
Um, so we need to provide a telside token.
I go run. Uh, we're going to get a panic when we try to run this.
Um so Telide token create a new bot using a botfather to get a token in the format of blah blah blah.
I wonder if I have I think I have this this token and I'm not allowed to share it with any of you.
Yeah, I do have this token.
My typing speed. Well, you've debated me now.
I have been debated into typing. Showing my typing speed. Let's go.
I will show you.
Oh crap.
Damn.
Whenever I make mistakes, it it's demonstrably slows me down like afterwards.
You see here mistake here and then immediate slowdown because I get careful. I get too careful in my brain.
Is that cargo cache or is your PC just super fast? My PC is fast. Yeah, that was the first compile. Um where is it? Uh 21 seconds in dev mode. typing simulator stream.
Good typing speed is a myth created by the bergies. The how do you pronounce that word? The I don't know. Bro is writing with his mind. Yeah, it's Neuralink.
If using Mac OS for cash, it's slow.
Linux has better file system. Oh, you're at 70 words per minute.
Uh 70 words per minute. How to reach my level? I don't know. I think to be honest, I've always typed this fast um since I was since I was young and I didn't really do anything deliberate to to increase my typing speed.
So, I don't even know like how I tried learning uh D'vorak and I tried learning a split keyboard and I couldn't I couldn't get past 40 words per minute.
Uh, okay. So, I need this token. I have this token. I do have this token somewhere.
Let me just uh SSH onto this machine and I'm going to crap um patterns file. Okay, patterns and file.
So we want a a e Q star and then that and [snorts] like everywhere I guess recursively but I don't want to show this necessarily. So let me pop this out.
Unsloth 3.6 Quen 2 27 billion. That would be too big. What's my setup? Which setup? My my PC.
You can go to uh my website visarchus.dev and that that information is here.
Bisc Bourgeois. Okay, going back to programming for a moment.
Is technical docs enough to build something or what else is needed?
Technical documentation when you're learning a new programming language.
Which technical documentation are you referring to?
All right. So what's actually going to be faster is if I don't look for this token, but I instead just uh provide it to this directory somehow. Okay.
I'm going to have to type it out.
which is super annoying. But you all can listen to my key taps.
Okay, looks correct.
So then um All right.
Now, let's try cargo run API side an error from the update listener.
API terminated by other get updates. I'm not sure what that means.
My tree sitter is all busted.
Why don't I use VLM? Apparently VLM is like way more serious than what I need.
Why do I have orange specs? It's for the blue light. Filters are blue light.
Argentina or or what are these?
Or Spain? I have no idea. I'm just going to say Argentina.
So, where I wonder where it got up to exactly. It was like This is the first thing we do with telkide.
And it cra it it died here somewhere.
Starting ripple. It didn't even say starting ripple, did it? Oh, that's not good.
It didn't even say starting ripple.
Yeah, it it almost logged nothing.
Oh, I have to rotate it.
All right.
Oh, actually that's good because uh this bot's called like Hermes bot and that's just not true. So, what am I going to call it? Son of Jet.
Username must end in okay.
Son of Jetbot.
All right, we got a new token.
>> [clears throat] >> check Lama then. I've got Alama. I've tried Olama. I I think it's fine, but some other people were saying that um Llama CPP is better.
I'm not having a trouble trouble running the model. I think it's [clears throat] more the um the harness is not like talking to it properly or whatever.
Like, however PI agent is set up with it, probably not the best.
Um, after this stream.
This is probably going to be the last stream that I do for LLM specifically, uh, cuz I don't think it's that interesting. Probably the next stream is going to be, uh, trying to make a game or something like that.
>> [snorts] >> Okay, so another instance of the bot is active.
That would mean that I'm running it on on um Why is this verbose?
Maybe it's being run here.
Have I tried open code? Uh, I think I tried it like once, but I haven't tried Kimmy or anything like that.
Yeah, Hermes is running.
Let's go ahead and kill this.
See if it uh comes back.
Yeah. So it is coming back.
What is running this?
Well, I checked system D and I couldn't find anything for Hermes.
Let's see.
Hermes, can I uninstall from here? Probably not.
I can uninstall from here. Yeah, cool.
Cuz I I don't want to use homies actually.
We'll uninstall you use Gemini. That's an unusual choice. While most coders I see in YouTube or personally know use Claude, I actually use Claude most of the week.
Um I I have a subscription with uh with Google AI and I don't have a personal subscription with say Claude. I think I can use codecs as well, but in general, I I really don't see I don't really experience a difference between these different models.
Um, what's I doing?
Okay. Yeah. So now when I run this again, it should be able to use my old bot which was compromised terminated by other. Really?
Is it still running?
No.
Hermes. What's Hermes? It is uh I don't even know how to describe it to be honest. It's an agent.
an agent with persistent memory.
But the problem with using Hermes with a small Jetson is that I don't have a big enough context window to to use it, I'm pretty sure.
So, I'd need more hardware probably.
or I'd need to hook it up to like um an actual um Frontier Lab or something like that. You use GPT 5.5 and Sonnet or Opus. Opus is the best. And with Caveman, it saves a lot of tokens.
Hermes site looks like an evil organization.
Um, Fable is available on the 20 USD plan. I've So, I've used Fable. I used uh Opus 4.8. I've used Sonnet 4.6. I've used Sonnet 5 and um yeah I like between it and anti-gravity not a huge difference for me.
Um, caveman. That's fine.
How do you actually set this up?
It's a skill. Uh-huh. Okay.
skill plugin builder investigator reviewer.
Okay, that's cool.
Skills cave crew. Cave commit. Cave compress.
skill.md [snorts] respond tur like smart caveman. Mhm.
Description.
This doesn't get included in the context, does it?
Just add the plugin include code caveman and then run caveman. Uhhuh.
Yeah.
Okay.
Can I set up plugins like that here?
skills.
Create new skills workspace global and shared.
Right.
Okay.
Any agent can support skills anti-gravity CLI skills. So I have none right now.
skill name and then skill.mdents/skillskll nameskll Yeah, I'm guessing that the uh I'm guessing that Claude code is better in terms of allowing uh customization of this stuff.
Okay.
Model [snorts] will also be focused on cave manning.
Yeah. Not sure how to benchmark that. I am. So, I'm still not sure why it's uh still not sure why it is unable to run the Telegram bot.
Goth Slayer 69 LLM question mark.
What's the question?
Oh, I just realized.
Huh?
Tasks. Uh, kill task.
All right. Now I should be able to run it.
And what was it? um test. I'm going to send a test message.
It's responded with thinking and then it gave back a response saying fetch page arguments URL and then google.com.
So, I'm going to say I never asked you to do that.
And it is now thinking.
And it says, I apologize. I don't have any specific instructions to fetch a web page if you haven't asked me to do so. I need a clear request to know which website or URL you'd like me to retrieve. Okay. [laughter] All right. So, it's basically it's kind of working. Let me let me try to actually fetch a page.
Uh fetch uh my blog.
Let's see if we can actually do that.
But I don't think it has logging.
Uhhuh.
So it literally just sent back the tool call kind of JSON message.
What did you miss in these past 45 minutes? Not a lot to be honest.
I think it's trying to call tools, but it thinks that I'm the thing it should be sending the tool call to.
Um, so I ran it successfully.
I sent messages. I asked it to fetch this.
It sent me back a JSON payload in the message which looked like it was trying to do some kind of tool call.
Kind of weird.
I will try to show you exactly what it sent me.
Let me just uh do a little sneaky sneaky down here so I don't expose all of my um photos to the world.
Uh, what's Gemini saying?
Okay. When you ask the bot to fetch a page, the rig framework registers the fetch page tool.
Uh, the interaction should go like this.
The client sends the prompt plus tool definition to the LLM server. The model decides to use the tool.
Why isn't this picture uploading, by the way? Super annoying.
user message llm tool call llm summary back to user.
Uhhuh.
I think Google updated their photo app and it's being an absolute [ __ ] Really, really disappointing.
Like where um literally where is my screenshot that I just took? This is crazy.
Absolutely insane.
I can't believe I have to do it like this, but I think I have to share it with myself via email because it's not backing up.
This is so I'm feeling the rage. I'm feeling a little bit of rage to be honest.
Um, I want to share a screenshot from my phone.
I just Okay. Uh, I guess I'll save it to Google Keep and then uh somehow I can access it from there.
Actually insane.
[snorts] Someone over at Google decided that this is a good user experience.
Keep notes. Save.
All right.
Technically, it should be there now.
Very good.
Wow. [sighs] Wow. Wow. Wow. Wow. Wow.
I can't believe how hard that was. But basically, this is what I sent on Telegram.
This is how the model responded to me.
Send me back a tool tool kind of call.
>> [snorts] [clears throat] >> Lameo, wow, congrats on 100K. Thank you very much. Uh, let me see.
Maybe fix prompt. System prompt. Yeah, the system prompt needs uh improvement.
Use local send from my Google pixel.
Does that work?
Is learning to become a data scientist much worth it than DevOps and QA? It depends on what you like to do. Do you like data science?
It's not looping back. It's doing this.
User message to LM tool call return to user. Uhhuh.
The client sends a prompt plus tool definition to the LM server.
Oh, local send. Okay.
How much of the setting up LLM work is done? Um, I've got llama CP. I've got an inference thing running somewhere. I'm just trying to set up this to talk to it.
Uh, probably better to wait for a summary if you if you're trying to catch up.
When you ask the bot to fetch the page, it registers that. Okay. Client sends a prompt. Tool definition.
Model decides to use a tool.
Preamble tells the model it can fetch.
If you're hosting it, use the Ginger flag that allows the server to use the model's built-in Ginger template to format and pass tool calls.
What?
Use a tool capable model.
Well, I'm using this one [snorts] to confirm what's going on.
Here's the combo image.
Uh, can I paste the image?
I can't paste images into this.
Let me look at the Let me look at the picture again.
Are you local local or remote local?
Local local.
Can I ask on Telegram to list all the tools? Yeah, let's see.
List all tools.
Oh, that's weird. It just gave me back a JSON which with a key and value name and list tools.
It's It seems confused.
I think that's the part where you intercept the tool call in the Rust code and then use an example.
Well, the code here, if we just look at main again, uh I'm going to open this in a separate tab.
Main. So, we have this tool which uh should intercept somehow. I don't know. Not sure exactly.
This is the web fetch tool.
Rust in the big 26. What do you mean by that?
C is 100% better.
[laughter] Are you trolling?
There's no way you think C is better than Rust.
That's crazy.
Okay, so that is how Quen outputs a tool call. So the the model is behaving correctly, but the server isn't.
It needs the ginger option really.
Is that really the fix?
That's cool if it is.
Uh, what we So, we literally just put Ginger here and that's it.
Running on a home lab. Yes.
I have a system build for this exact scenario. Web fetch file creation. Can you ask LLM to have a look what I said?
User a loop.
Yeah, I think we'll get to fixing that.
You have already have that implemented.
What do I have implemented? Sorry.
Would I review your resume? Yeah, send it to my email. Rust isn't better than C though. It is better. It's way better.
You can maintain CC code with Claude.
I mean, Claude will certainly help you maintain CC code. That doesn't mean that you're going to be absent of uh foot guns.
You got to 5070.
How much VRAM does a 5070 have?
Yeah. Yeah. So, I already added the tool call apparently.
Prompt it until it works. Huh.
That's crazy. 12 gig. I see.
All right.
Uh chuck a little damon reload and then uh we're going to give a little restart to the llama server.
Um, and then maybe I'll like restart the the son of jet.
Let's see. Is this now good?
Uh, it has it has the ginger thing. It started. It looks okay.
Gonna run this see with super armor.
I don't know. Like you can bolt on protections like they have this um fill C thingy.
That's fine, but I still think Rust is just better.
uh fetch.
Let's see if it will fetch my website.
It's thinking and it's giving me back the uh the tool called text again.
um tells the server to load the model's native Ginger template, format request correctly, and pass the models JSON output did all that. It is still giving me back the JSON. I think it is missing something to tell it to give that to uh oh so let me think about this right we're running uh this rust program okay which is has created the agent and given it a tool but but And yeah, so we're receiving the telegram messages here, right?
Okay. I think the problem is that the Rust program is uh receiving the JSON payload from the model, but then it's not doing the tool call and is instead just sending that tool call as a text message back to my telegram chat.
I think that's the problem, right?
So, this literally son of jet should be picking up that tool call and uh handling it here and then giving information back to the model and then it should then send a message saying this is what was there.
What's Gemini think here?
That's exactly what's happening. The reason the uh rig is sending it back as text is because of where the JSON is located in the API response. So it expects the server to return the tool call inside tool calls message roll assistant tool calls function. Oh, really?
And what my thing is currently doing is something like this.
Okay. Okay, cool. But it is actually running with the ginger option. So the tool call should go back to LLM.
Yeah, that makes sense.
Since I'm running with the ginger flag, it should theoretically be able to format the tools and pass the output.
However, it's still outputting raw JSON.
We need to look at what the Okay, when you send a message, what does the console output on the Okay, let's take a look. See, then we should be able to see like some uh some stuff, right?
Uh Walmart server.
Oh, what is all this?
This is indecipherable.
Oh, wait.
Hold on. What did it say? Context looking token.
This was not a control type. Okay, I think that's fine. This is it talking to the actual bot.
Can I get no time stamp? Is that a thing?
No.
Cool. shades. Would it be easier to set rigs logging to a boat? Yeah, probably.
I really need to fix my um my tree sitter. Let me see.
Lazy clean.
No.
L.
No, that's not it. Wait, where is my The model's built-in chat template in the GGUF file lacks tool support or the llama server fails to match the model's text formatting support. To make your bot fully resilient, I've implemented a client side interceptor loop. Oh. Oh, no.
Oh no.
Query the agent if some tool call extract JSON tool call.
Let me see this.
Ew.
What?
This is terrible.
Give it a spin.
Changes it. Blah blah blah.
Let me just see if that actually works.
We'll see.
Copy my last message. Send it again.
Thinking bching page.
Taking a while to fetch the page.
Okay, it has fetched the page and turned it from uh HTML into text.
That's pretty cool.
Don't know if I can show you all going to do the I'm going to do the dumb thing and show you uh my actual phone because technology sucks. But it says the web page fetched appears to be the personal website of Vasilio Sarakus.
Page contains information about the website owner, his blog, and various details, which is true. That is true.
So that's kind of cool.
that it it it it worked, but um the fact that I had to like do this little little interception thing kind of weird.
Not sure how I feel about that.
You got the problem.
Did you send something?
I think maybe my messages haven't been coming through for a bit.
Last message I got was, "I got the problem. Now let me clarify it.
The user sends message LLM then says here's the payload. Yeah, I understand that payload should be handled by the system to call the tool. Completely correct. Yes.
The system is returning the LLM response which is the tool called payload.
Correct. Yep. Yep. You're 100% correct.
Downloading anti-gravity now.
>> [snorts] >> Um, yeah, I think there might be something wrong with my messages right now.
I just let me see test Maybe I lost connectivity to uh to YouTube chat or something. Maybe I'm getting rate limited.
No messages coming through.
Oh, it's so janky.
Yeah, I I uh I vibe coded this one. I've got a video about this one. Um it is uh this video here, building a small tool in Rust.
Uh basically I made a little um a little chat application which uh it cues up all the messages from uh YouTube and also Twitch chat.
Um so if I if I pop over to uh Twitch, I'm not streaming on Twitch, but my chat if I send message in here, it should show up here.
But also YouTube messages should show up here. But the but YouTube is a bit different. On Twitch, you just connect to an IRC server and um and you can get the messages. But YouTube, you have to hit an API to get messages. So, it's a bit crappy. It can extend if we add L1 there. just to respond back to users.
See, yeah, these chats aren't aren't showing up at all. So, I'm going to just uh just going to pop this chat out so I can still see messages.
Just joined. Where's my MCP setup? Uh, I'm not using an MCP for it at the moment.
It can extend if we add LLM there just to respond back to users or just comment funny and chat with users. I would like to set up uh a chat bot particularly for answering common questions.
Like right now we have 50 viewers, but usually when there's like 200 people, I get asked what operating system I'm running at least, I don't know, 20 times over the stream.
So, it' be good to have a bot to answer those.
and also about my keyboard and that kind of stuff.
Yeah, there's no there's no commands at the moment.
There's actually no bots in here at all.
Um, so yeah, maybe what's this? Ga should be set up auto reply to should I quit uni?
What if grease monkey watches chat and posts to me slashssage on dom updates?
Yeah, I don't know. Uh, I think I should be able to um get a bot set up. I just haven't put any time into that. Okay. So, the last thing here was uh um also I I don't like this.
I was doing um so basically if the message starts and ends with a uh curly brace, it tries to serialize it.
It's kind of a weird way to to do this, right? Because I could just as easily just do this. No.
Isn't this literally the same thing?
Like all all of that is super unnecessary.
What were were I in was I in France San Francisco? No. Uh I'm in Sydney, Australia.
I was at Atlassian between 2018 and this year.
You've built many projects with Vibe Coding, Speedwatch, Dev Thread, currently working on local learning.
Cool.
feel free to um like send over your your GitHub or post your GitHub in the chat.
What model am I setting up right now?
We're working with Quen 2.5 with 3 billion parameters. So, it's pretty basic, pretty small, but it's it's it's snappy enough.
Like if I open pi and I just say something, it is uh it's not bad.
Like I'm not sitting here waiting 10 seconds for a response. So it's not terrible.
How do parameters matter exactly? I personally don't know. All I know right now is that if there's lots of parameters, you need generally more RAM to handle it.
I use that for an in browser model to create reports on analytics data.
Which LLM model is this? It's Quen 2.5.
with three billion parameters.
The only thing or bottleneck is RAM.
Well, it's an interesting one, right?
Because the Jetson has very small amount of RAM, but it has a processor which should be very good.
3.66 Six six billion parameters can fit on on this machine really on my Jetson.
Well, I could give it a try.
You ran a model of around 4 billion parameters using OAMA. The response was slow. Your GPU is 30 54 gig. That's quite um that's [clears throat] quite a low amount of VRAM.
Try a bigger model with offloading.
Can you explain what offloading is?
[cough and clears throat] Let's see. Um so antigraph is saying when rig constructs API request to alarm server it attaches the web fetch tool.
If alarm server natively supported the tools API the server would inject the tool definition.
So is it saying that there's some interaction between uh rig and the inference server?
If you want native tool calling, you need to align llama server with the openi tool spec.
So run the server with the verbose flag.
When I message the bot, check the console output.
Okay.
Please hold while while we set this up.
Um, so we fixed the issue. The issue was that when when I was So I I currently can talk to my Jetson via Telegram, which is cool.
But when I ask it to use a tool call, it is um just giving back the tool call message as JSON to my Telegram chat.
We fixed that by intercepting the JSON and calling the tool ourself, but it needs to be better.
The llama server can't access tools.
Maybe let's chuck this on verbose mode.
Let's see what happens.
So I'm using some rig lib in rust with llama CPP.
Correct.
Uh let's see.
Uh what do I ask the model? I guess I'll just ask it to fetch the same page again.
and we'll just see what happens in the logs.
Error from LLM agent exceeded context uh size error requested 4,300 tokens and the context size is 4,96.
>> [gasps] >> You can buy two to four H100.
Sure.
Oh, okay. So, it's uh 40k per card. H wait. Fully bundled 8GPU server is uh 300k.
Yeah. Yeah. [snorts] It probably sounds like a jet engine taking off when you put this in your computer.
In your uh in your house and switch it on.
Uh okay. So it fetched Yeah. So it fetched the page, right?
JSON number JSON blah blah blah. What's all this?
GPU effects reasoning. I don't know.
Past message roll assistant content.
I think this is the key here.
Update llama server. I think it's I think llama server is okay.
Run this. Uh, change the bot lama.
Try Olama.
Okay. So, what's the difference between like why should I use Olama versus Llama CV?
So, Llama CPP we have maximum control over parameters like sampling, context, batching, GPU layers. It has the best performance.
Uh, so yeah, so I compiled Llama CPP for CUDA.
So that's cool. Easy model switching, built-in REST API, AI compatible, blah blah blah blah blah.
Am I trying to understand or just trying to run it? Uh, bit of both.
[clears throat] differential. Wait, hold on.
Common chat templates apply ginger.
It's saying update the preamble.
name arguments sequence template applied generated parser.
Wow, crazy.
Uh launching slot.
Wait, why is it listing out my uh what's this? Neogit.
Why has it got these? Do I have that on my site?
Am I list Am I have I got Neo get here?
Oh, I do. Yeah. Okay.
All right.
Uh, I'm going to take the verbos off. I'm also going to just quickly look at, okay, we'll just restart that.
Um, llama. So, here's all the commands we have.
Uh, llama server llama simple.
Um, what else do we have? GGUF.
I don't think I need to run anything different.
Just need to um anyway uh what did it change?
It changed something here. The preamble.
When the user asks about a website or requests you to look up the URL, um, you must call the fetch page tool by outputting a JSON block wrapped in tool called tags.
What? Dude, that doesn't seem that does not seem like what we want.
Right. That seems like the exact wrong thing to do here.
Right.
So the last fetch that I got it to do, it outputed an error saying that um that uh we exceeded the context window.
Okay, cool. Why does it say starting Pi agentbot?
Wait, why is it Hold on.
Starting pi agent bot.
There's no pi involved here. Anyway, let's uh send a message.
Receive message.
Querying agent.
No tool call detected in response.
So, we just sent back the JSON again.
Did I code something in Rust that you're debugging right now? We vibe coded it.
You want to learn and build page of duty and ingress type stuff in your TypeScript mono project.
You're going with GPC and go I cannot comment on whether or not that's correct.
That's a design choice.
Is this something I should not leak?
Is this private information?
I think it's just the chat ID. I don't think it's my user ID.
Am I Aussie? Yes.
You love the videos. I appreciate it.
Yes.
So for some reason it is not picking up this tool call thing. Name arguments names a string and arguments is a fetch aug which has a URL which is a string.
Yeah, something is not right.
this got laid off a month ago. Your video is really inspiring to keep going. Cool.
How's uh like what's going on for you right now?
Uh, let me see. Let me see this. Ryan, rig core. Does it have examples?
Simple example. We got client and then agent and then uh prompt.
I don't see it ever actually prompting, do I?
What would I choose?
I mean, you can't really go wrong with gRPC and go to be honest.
Um, chat ID, user text, received text, and then clearing conversation history. We don't care about that.
typing message loop.
You're an S sur in Europe applying nearly everywhere and almost never hearing anything back.
You submitted your CV at midnight on a Thursday and 15 minutes later got rejected. That to me tells me that something might be wrong with your resume because basically they So what I would suggest in terms of your CV is set up your CV how you like it and then get Claude to re to uh reword it. Keep the same meaning but reward it because everyone's using LLM tools to analyze résumés.
and LLMs prefer LLM written CVs.
Have I worked with Go? Yes, a little bit. Okay, so here it is. Agent.
Agent chat mutably borrows history and appends the turns to it. What?
Oh, our current uh Telegram bot agent rig is having issues.
Uh when the model sends us a tool call, it is not being recognized.
I am not sure if we should be using the prompt instead of chat with the history.
That might maybe another matter.
Job market is [ __ ] Yes, job market is [ __ ] You know what I think?
There's been so many people laid off, right?
Thousands.
I feel like all the people that were laid off should band together and make a company that sells something to the companies that they were laid off from.
If you want to learn gc and go, where should you start?
Do you know Golang currently?
Do not tra try same thing CV experiment on that constantly.
Same issue on putting in CVS and had to rewrite. I just use chbt and tell it to customize for the job and also ATS.
That's cool.
Yeah, you should probably prompt it for ATS. I did one uh uh redesign. So, here was my Wait, here's my old one. This is my old CV, right?
And then, uh this is the one that Claude made.
Why is it taking so long? All right.
Yeah.
So, it it rewarded stuff.
Split this.
Uh expand.
Claude on the left. me on the on the right.
Uh I I like Claude's design way better.
Leaked the number. The number is public on my uh website.
It's not a big deal.
I don't answer my phone anyway.
But yeah, so this I probably should get it to rewrite it again for uh uh ATS.
Make an enhanced version for ATS.
Uh um but retain original semantics.
Don't embellish me.
It's going to be a problem when you have a million subs in one month in a few months. Try with Fable.
Uh I don't think I have Fable.
Yeah, it's pro. It's in the pro plan.
You love what China is doing with open source models. Yeah, I want to try Kimmy K3 and and these other um I want to try them, but I don't I think I need like open router and uh all this kind of stuff.
I haven't looked into doing that stuff.
Okay. So, what's it said? [sighs and gasps] In your main, the inline tool extraction was relying on SR.
It expects the entire string to be a perfectly formatted JSON object.
However, LLM's rarely output just JSON.
Uh-uh.
Oh, that's why it was doing that stuff.
Oh, man.
Yeah.
And then uh am I single? No, I'm married.
Some ATS's don't really like grids or multiple columns.
Really, what a crappy tool.
Um important note on rig framework and tool calling. Using agent chat is entirely correct for keeping conversational history. However, you have a manual loop in your code to execute tool calls. The rig framework already handles tool execution and editing loops.
The reason rig's native tool calling isn't triggering automatically in your project is because your initialization uses the completions API.
V1 completions is a legacy text continuation API and does not support native tool calling. really.
Okay, let's try that then.
So, can we just like uh pop that away and then >> [snorts] >> I3 is better.
I3 K3 just get open code CLI and make an open router account and set up open code with open router. All right, I'll do that later.
You have to pay uh it's like API costs for um open router. Is that right?
Open router.
They're like Open Router is the one that runs the inference. Is that correct?
$5 a minute.
Uh, okay. Kim K3 like this one.
$3 per 100 mil of input tokens.
$5 minimum. I see.
Historically, they didn't like it, but as they advanced, some now got more extraction capabilities than others.
I don't know why it logged me in straight away. I actually didn't want that.
[cough and clears throat] Yeah, this seems like the the right page for pricing.
Let's look at uh let's look at like set five for example. Claude Sonet 5 and 10. Do you think that's right per 1 million?
Is Kimmy actually more expensive than Sonnet or should I more compare it against like Fable?
Oh, 10 and 50 and 50. So, it's like three it's more it's more than three times cheaper than Fable.
But then if we look at uh 56 soul, right?
Okay. So Kimmy is half as expensive as Soul.
K3 is 70 terabytes. I don't think you're running that locally. No. No. Absolutely not. I'm not running that locally.
Let's look at the one that I'm using locally, which is 2.5 um 4 billion instruct.
It's not even here, right? Let's just look at 7 billion.
Wow, that's cheap.
It's like a 100 times It's like uh 10 times cheaper than No, it's a 100 times cheaper than Sonnet.
That's crazy.
Are you telling me like think about just think about the the like the output that you could pro imagine you had a 100 attempts to get the prompt right compared to like Sonnet.
You know what I mean?
Just think about that difference. You could run Sonet once or you could run this shitty model a hundred times.
Do you know what I mean?
You're paying with your time.
Yeah. If you're if you're running a task that is you need it to be responsive now, then it would suck.
But if you're running something in the background like a research task or let's say you want an agent in your in your application.
Let let me let's see haiku.
Haiku one and five.
See that's still um what is it?
100 lums of [ __ ] is still a hill of [ __ ] But if your task was basic, you don't need you don't need haiku. You don't even you don't need like imagine that. Let's see what's a task that you could use.
Um, okay. So, let's say you had uh like your bank transactions and you're like, I want to categorize all my bank transactions and I want to figure out which one is like um tax deductible or not. Do you get Haiku to do that for $5 per 1 million output and $1 per million uh input or do you go 50 times cheaper and do that instead?
Cuz it's some task that you don't need immediately done. It just has to get done at some point and then sent off to your accountant or something. Do you know what I mean?
I think there is a place for quote unquote bad like terrible models. Do you know what I mean?
I think there's a place especially if you're maybe you're running a SAS or a business that is integrating LLMs into your actual application so that when users are doing stuff there's all these agents being called and you have these little um things maybe you're using like Um, Telegram linking is free. Yes.
Yeah.
I know Cloudflare, they have like Cloudflare workers and you can integrate AI with Cloudflare workers and and they actually h they're running GPUs. So, you can use local um like open- source models for way cheaper. So, if you're a business and you're running Cloudflare, you know, Cloudflare is really cheap for a small business. And um not sponsored, by the way. And uh you know, you could decide I'm not going to use anthropics models. I'm going to use some open source model to do these little tasks.
They chain good models expensive at the top doing the complex and subm models for the delegated tasks right yeah it's just that fine balance I guess when you really figure out like exactly the task and it's a predictable task at that point yeah It's hard to make this call where it's like I just I use a model to do it or I use a really good model to make a deterministic tool to do that instead. I don't know.
Yeah, [sighs and gasps] it's tough.
Um anyway, back to looking at what this rig's doing. I think we need to cancel this and rerun it from someone who works on AI native HR techch. Yes, we run a lot of crappy small models for some granular things and it worked perfectly.
H yeah, that's what I thought.
Trying to work out that now too.
All right. So, fetching page. So, I just messaged the bot on Telegram, told it to fetch a page, and then it uh exceeded the available context size.
It works without the Wait, did we querying successfully detected in line.
Right. Right. Right.
Um Have I tried running a small model on my phone? Okay.
What's the end goal of this bot?
Um, that's a good question. So, my the the picture I had in my mind was that I can talk to the bot from Telegram.
There was some like maybe like analyzing my bank statements was one thing I kind of wanted to do. I'd also like it to like um handle some of my um maybe my emails, maybe my um LinkedIn connections and [ __ ] like that.
I don't know. I have to think of use cases, but right now I'm I'm just trying to set up this the capability.
Usually has a debut trim about things you want and just go ahead and build it before thinking about what the actual end goal is.
Uh what was it? context.
Yeah, I don't know why. Um, there was a suggestion somewhere to limit the context size to 4,000.
You might want to bake prompt injection prevention in.
Yeah, that's true. Otherwise, someone could hack me through my LinkedIn. That would be horrific.
Do I only use near? Yes.
Use Corsair if you want to let AI handle stuff.
Yeah, there's stuff like uh Hermes.
What's What's this Corsair? Is that an agent?
Cool.
What are their integrations though?
>> [clears throat] >> It doesn't seem to have um like local agents, hub, concepts, plugins.
Okay, cool. LinkedIn. Nice. Nice.
Does LinkedIn have an API?
Do you have to pay for it?
Verified. Sign in.
Marketing. I don't care about marketing.
I don't care about talent.
I care about messaging and accepting connections, consumer profile. I don't care about that.
Real man use the gooey.
What if it's not a man using the gooey?
It's a robot.
you abuse URL params plugins.
Yeah, it doesn't look like they have like a messaging API or connections API.
That'll be cool.
Toxide's running fine.
Got rig core running.
mostly class classifiers ML model.
So you're using uh LLMs to do class classifying.
What's it say?
[clears throat] Ma two main reasons for a context window filling up. large webpage content persistent history buildup. Uh-huh.
Let's try clearing the the history because I think it's including all the previous messages.
See, um, now when I do the same, uh, prompt to fetch the page, let's see if it comes back with the page.
No, it's still hit a context uh, size problem.
history cleared. It ran into the same context size issue. However, I was able to get this printing summary of the page previously before we made changes to message is intercepted.
When I get GPT to scan public LinkedIn, it says you it says it uses API H.
[clears throat] Okay.
But are you scanning people's posts or are you um is it like [clears throat] let me see. Can I use uh an API to accept or deny connection requests on LinkedIn?
[clears throat] Are there any good use cases for local LMS? I guess you have a MacBook.
Um, the models are incapable.
It depends on what you're trying to do.
I'm not trying to do anything extremely advanced. I'm not trying to like code with a local LLM. I'm just trying to get it to do dumb small tasks like uh I don't know, maybe reading documents that I have or something like that.
So you can post content and pull profile data, but you cannot do oh you need you need to have you need to be an enterprise partner for messages and whatnot [clears throat] local before we fix This the script often failed to intercept the text base.
When it failed, it simply returned the raw JSON text to Telegram and exited the loop.
Once we fixed the extraction reax, the script successfully intercepted the tool call. However, it then passed the result back to the LM as a standard user message saying here is the content.
They get confused when tool results are fed back to them as if regular user messages. Instead of summarizing the text, the model likely thought, "Oh, the user wants me to fetch a page and outputed another JSON tool call." I don't think that's what happened. In each of the five loops, your script appended 2,000 characters of web page to the history. Really?
Huh?
No way.
What am I role playinging for? Me?
I'm not role playing.
What am I using for web search? Just like a HTTP client.
It panicked.
That's crazy.
Nice fix, I guess.
simple writing or role playing.
[gasps] What the [ __ ] is it doing? It's crazy.
Why am I using How come I'm using Gemini? just because I have a subscription for it. That's all.
Okay.
So, uh this is basically what it sent back to me.
this message right here. So, we fetched the page and it uh yeah, it kind of heavily summarized uh what's there.
I guess it's summarizing that because the uh this agent is um converting the HTML to plain text.
So, this would have been all it received.
You enjoy my stream. Thank you.
How would I measure token generation speed? I don't know.
Okay. Okay. So, what's a let's say um let's say this one uh fetch um docs.rs slash ririg core.
Yeah, let's just give it that.
fetching page and it's responded saying it's a rust library provides various functionality for handling processing data passing and processing of text data.
Wait what?
This is so wrong.
It's owned by CV Alclair.
No. Oh, it is. Okay, that's true.
But it completely missed that this library is about LLMs.
What?
All right. Well, at least it's doing something. I need to do a full better system prompt can uh improve response.
Yeah.
So the code that is made here is kind of gnarly. So here it says this is the preamble. You're a helpful AI assistant.
You have the ability to fetch web pages using the fetch page tool.
I don't like that I have to even explicitly tell it that that's the case.
And then um this part here where we're we're doing this max loop edit text message querying agent for chat current prompt and history receive response if It's a tool call.
What if I just um What if it's not a tool call?
Wait, why can't I comment that nil?
If I completely remove that bit, what's it going to do?
Yeah. So it it just prints it back. So I don't I don't understand like why am I why what is this trait for? Like [clears throat] I want to see an example.
[cough and clears throat] We're it's it's implementing tool and then it's providing the tools to the agent.
And then prompt how much would blah blah blah prompt. And then there's no explicit handling of uh of the response agent. And inside agent we have this agent strct which has I think uh preamble hooks model name client and then build [clears throat] preamble tool. will build.
Where is the tools here?
Is it part of another trait?
We've got agent builder, right?
So uh this build why is there three with tool server handle?
Basically we get an agent back and the agent has a chat trait which is what we're using. Send a prompt with optional chat history to the underlying completion model.
If the response is a message, it's returned as a string. If it's a tool call, then the tool is called and the result is returned as a string.
So, this should be handling the tool call.
Got chat here.
Is this not implemented?
Is there like okay uh it would be uh it would be in the agent. So agent what are the in here agent it implements chat let's see the source code here I imagine that the trait implementation is somewhere in here input completion input prompt prompt request from agent typed prompt tests.
Okay, don't care about tests.
Imple chat for agent.
Let response equals prompt request from agent.history dot extended details from agent prompt.
So prompt request from agent basically is the thing.
But where is that from?
Streaming prompt request. Agent prompt request agent prompt request. Is it here?
No.
Do a little search if we can. Prompt request.
Uhhuh.
You going to head off? No worries.
Thanks for chilling.
Check rig providers specific crates.
Maybe there's an impul there. I think we're on the right track right now.
So here we have uh prompt request from agent got the agent passed in reference and then there's a into message and then it's going to give back what prompt request.
Oh okay.
Um, back.
[ __ ] I lost track of where I was.
I was looking in the source of this one.
We had prompt request from agent and this was in the implementation for chat. Then you've got dohistory.extended details.awwait.
But this is just a strct, right?
extended details just adds more to it.
But then uh there's something async in here. So send response output run a run send Hello from Wellington. Cool. Welcome in.
Okay. Prompt response then.
Whoops.
Yeah. Is beef Wellington from Wellington?
Output usage completion calls messages.
So what I'm missing is the part where it says how it works.
So, we're going to look in this actual repo.
Explain how the agent detects when a me a prompt response contains a tool call.
Yeah, I think that's what we want to know.
Oh, Wellington, New Zealand.
Mhm.
I'm feeling like I'm feeling like taking a rest. How How long I've been streaming?
Two and a half hours.
Maybe I play some games instead at the moment.
Do I watch FIFA? No, I don't. I don't really follow any sports.
Sleep is also an option.
Extract JSON tool call is never called.
That's fine.
And we can keep going on this for a little while and and then maybe play like uh try to get a win in Noer or something like that.
Do I like techno? Yes, I do actually.
What kind of techno? H What are my options? [snorts] I couldn't really name any artists to be honest.
I do have some music on that I've generated myself.
Maybe I should link Sunno on my uh page dark wave. Sounds cool to me.
I think I should link this on my um on my page maybe.
Hypnotic hard techno. I think hard techno would would be probably what I like.
Um hard techno. It depends on what mood I want to be in. Um, what's something I've listened to recently?
I don't know. My relationship to music is that it's like um a tool that puts me into a particular state of mind um for an occasion like if I'm trying to work concentrate or if I'm trying if I'm driving or if I I'm at the gym.
There's a lot of eye music at the moment. Yeah, apparently people don't like it very much.
Um, but I've I quite like it.
I will actually my profile has a bit of weird stuff on it. Maybe uh it's not like exactly representative of what I would say I'd like, but I I like I guess I like every song that's on here.
Um, let me see.
There's like a decent amount on here.
Not everything is like public. I haven't published everything.
How do I add this to my profile? Give less mealis EP.
Is that Kavinsky Night Call. I am pretty sure that I know what that song is and I think I like it.
Customize.
I wonder if I should link to my suno or if I should just uh uh upload it to YouTube.
Maybe I should do that.
Maybe I should do that.
All right. I'm not going to link it there, but I'll link it in the chat.
if you guys want to go see some of that stuff. But then I've also got like music uh YouTube music.
Uh let's see. Let's look at my liked.
Is this even recent? This doesn't look This doesn't look accurate.
This doesn't look uh from the community. Quick picks. Where's like the ones I listen to?
Bonky kind of stuff. And do I have to switch account? Maybe Factorio.
No, this is wrong.
Oh, yeah. This is right here.
So, this would be some of the stuff that I listen to when I'm driving most of the time. I really like this one, Agrazek.
He's sick.
You like my tattoos?
Sometimes I forget that I have them.
Uh, what was one of his recent ones that he These are all the Sombra. Yeah, I like Sombra. It's pretty good.
upload and monetize my Yeah, maybe I got a bunch of stuff on So in my library that I haven't uh published.
Tons tons and tons.
I I got one to kind of sound like um uh Lana Del Rey.
This one here.
Uh I'll publish this one.
Let me see if I play it. Is it going to be too loud? What's this say?
The walls are whispers. They grow.
A tower of built on.
>> Would you know that that's that's AI?
>> It's like elevator music. [laughter] >> That's hilarious.
That's not elevator music.
>> I don't know. I think it's pretty good.
Yeah, that was AI.
Crazy, huh?
Crazy.
Anyhow, uh, wait, what's going on here?
[clears throat] Just 8 gig of V RAM that can barely run quantized models. I'm not using it for coding.
I'm not even going to attempt to use it for coding.
[clears throat] Crazy times. Yeah.
Uh the rig library is designed precisely to abstract that logic away. When you register a tool, it detects if the thing since you're connecting to a local model, this native tool execution relies on your local back end supporting it.
Okay. Okay. Well, then if that's the case, let's just jump on here real quick.
Uh, we have Llama CPP installed. Can you check if it supports the OpenAI native tool calling API? Was it structured tool calling API?
Why do you use anti-gravity not CCL codeex? just because I have a subscription with uh with Google AI.
>> [clears throat] >> Let's see on the actual uh this on the actual uh llama CPP GitHub does it explicitly say that it supports tool calling.
What's this?
Completions. That's cool.
Uh, tools down. What's this?
bindings.
Okay.
Infra models obtaining and quantizing models.
Enable specular decoding.
Have I tried soul and fable? I've tried fable. Haven't tried soul.
What is it trying to send? Uh, it's trying to send a message right to the API.
It has written a script for debugging the functions.
Yeah, I'm keen to try Soul. I do have Well, I'm not paying for um Open AI at the moment.
Thanks for the fight fight brain rock with worked examples. It helped me a lot when dealing with a lot of entities in an old codebase.
Cool.
What do you mean by entities?
What did I think about Fable? I thought it was way too expensive for uh what it was doing.
Like I would I would use Fable for a task and for that same task I could just use Opus and it was like 10 times cheaper.
So, I don't know. It just seems crazy to me.
You work in a SAP related codebase, huh?
Interesting.
[snorts] Um, does this support open AI structured tool call API?
Pretty much never use Fable unless I'm auditing a security fix or something. H.
Um, so this apparently natively supports um the tool calling thing. And apparently all all you need is this gingerbased chat template, but I don't know.
I'm going to let this uh complete its investigation and then uh I'm going to play some noter.
What's this saying in the rig architecture? Agent detects when a response contains a tool call. Here is the exact mechanism.
Assistant content ah detection in the agent loop.
When the model completes a turn, the agent collects these structured content blocks into a vector.
Once the agent finishes its run, has tool call. Any block matches block assistant called tool call.
Right.
Do I have a job yet? I do. Yeah, I've been working for like two months, but it's hush hush right now.
Hi, I just found this live. One question is local LLM's good for coding. You need a good m a big machine. Um, I would probably just pay for claude or chat GBT. Yeah, you'll spend more money trying to figure out how to get coding to work well locally. There's no point.
This here llama template analysis. Okay.
You used soul today and you created a full solution with agent integration for Azour. Took 30 minutes and only use a few% of my quota. Was testing it when I saw this stream pop. Cool.
Yeah, I suppose like I have a bit of a business idea that I want to maybe try and um so I could maybe get Fable or sold to make a proof of concept.
You have 64 gig of RAM and uh 5080.
H I suppose you could probably run a decent agent locally.
I'm not sure exactly uh how capable of a model you could use, but that seems like good enough specs to do that.
Okay. So, what's this?
This is running.
Tool choice required. We've got get current weather parameters properties roll.
That's interesting.
I close this one down.
Which one do I recommend? Like which model for coding? Well, all of the all of the rage right now is Kimmy K3, but you could probably, as far as I understand, um, you could get away with something like Quen 3.6 uh, coder Quen 3.6 six 27 billion parameters coda.
I suppose I'm sure someone else in the chat uh knows better than me to be honest.
I imagine that model would be pretty decent on that machine.
What is going on with this tool call?
Uh, I haven't used open code enough to really form an opinion on it.
Reasoning equals thinking with a [ __ ] ton of tokens. I hate that that's the best they can do.
Uh task like what's going on with this task?
Do I play Battlefield 6? I do not.
But I do like Battlefield.
Battlefield, I don't know if it's the same, but do they still have like roles in the game? Like you've got medics and the guy that drops ammo and stuff like that.
I like those kinds of uh role-based multiplayer FPS.
One of my favorite ones like that is uh Wolfenstein Enemy Territory.
I love that game.
Okay, so Llama CPP fully support structured tool calling.
Uh, cool. Cool.
Ever played HalfLife? Yes, HalfLife is an amazing game. Uh, like Halflife 1, amazing. Halfife 2, I can't remember. I don't think I ever finished it. And then HalfLife Alex was really cool.
Anyway, we're going to play Noa now because uh yeah, I'm going to play Noa. Going to win one match, one uh game, and then probably call it after that.
I'm not on Twitch. No, not yet. I um I think I might set up like reream at some point. I'm not sure.
The infinite generation loop problem is scary.
Last point in findings.
Tell me if this is too loud. I think it's not too loud.
Someone put Counter Strike on a browser.
Cool.
Terraria vibes. Yes, it does have Terraria vibes, but it's there's not really building in this.
There's you you make wands with spells, but uh that's about it.
Uh, yes. This is a rogike. Yep.
You're putting Terraria in the browser.
Interesting.
Twitch is better is what you have heard.
I was streaming on Twitch um earlier, but then all of my subscribers are here. So, I'm probably going to stream here for a bit and then I'm going to uh look into doing a double stream across both at some point.
What's the name of the game? No.
Holy moly, there is a worm.
Uhoh. Okay, since you already have subs here, YouTube works. Yeah, I actually asked people in a poll um where would they would they like to see games and where would they prefer to see them and 65% of people said uh yes and that they would prefer YouTube.
But yeah, I'll probably try to do both.
can make can you make decent coin on from Twitch for smaller accounts? I'm not sure.
I guess it depends on how your community goes with um like subscri subscriptions and donating bits and all that kind of stuff.
Nowadays, kick is better for streaming.
Really?
That's interesting.
even for like software type streams because I I don't think any of the software people are really on on kick.
want to avoid kick. It's shady.
Everyone on kick is just bots.
How to win in this game? There are actually several ways to win in this game.
Um, the main way that people win is they Well, it's kind of a it's kind of meant to be a surprise, huh? Let's see if we can win this one.
So, I got 300 gold right now. Isn't kick all the rage beta kids? I don't know.
Another one.
Why are there so many of these dudes?
Precisely why I want to be there. So that it's sane here them to be there. Ah okay.
What do you mean by their rage baiter kids? Like who are they? Who are they raing like on Tik Tok or something?
Came here for local LLM but ended up watching it game. It's Nita.
No.
[groaning] The ones that go around baiting people, they normally have security with them.
Oh, those uh I'm pretty sure those are just terrible human beings that probably don't deserve to live.
Big frog.
Oh, that's nice.
So, have you set up the local LLM?
Um, we've got it to a particular state.
Basically, I've confirmed that Llama CPP does have structured tool calling built in and I've got um Telegram set up so that uh I can message the bot, but there's a bit of mishandling going on with the the the way that the tool calling uh is being intercepted, it's kind of not working very well, but we'll fix that.
But yeah, given it a rest for the moment.
[clears throat] Um, let's see.
I'm going to regret that maybe Ow!
The power of the tablet. What can I say?
Which LLM you put? Uh the model I was using was uh Quen uh 2.5. Nothing nothing extreme.
Something basic.
The inference server was llama CBP and the harness or the agent I guess is uh I'm making a custom one that I can talk to from Telegram.
Okay, there's nothing here.
Well, so I could look around for like a nice wand since I still have a bit of health.
But this layout is kind of bad.
Um, can't really break through there. Okay.
Why can't I kick that open claw?
I don't know about open claw.
I tried Hermes agent. I didn't really like it.
I also I'm not sure how well those kinds of harnesses work or those those um agents work with a smaller machine and a smaller model.
Damn, dude. I hate these shotgun guys.
There you go.
Join late. What model did I run? I was running Quen 2.5.
Um, the model wasn't really the the focal point. It was more just trying to get it all working together.
Hermes, an open claw of shiny objects, unless you actually know what you're doing set up to the point where it's for you.
Yeah, that's why I want to use this rig thing. It's like I guess you could call it like a a micro framework rather than a fullyfledged model. Like I don't want to just strap together everything um like a bunch of stuff that I don't necessarily need.
Uh, very cool. I'm running one as well, although my 4 gig VRAM is only allowing me to do 12 billion. That's 12 billion even seems like a lot for 4 gig. I was doing three billion only. You recommended Kubernetes the hard way.
[laughter] That's funny.
Then you later found out that there is cube adm. Yeah, I got pretty frustrated, but it's all good.
I think um like if I got more serious into Kubernetes, I'd probably read that guide again.
But I'm I like to get a bit a bit pragmatic and scrappy at the start spells.
How long is a run usually? Probably half an hour or 45 minutes.
Depends how I go. Really?
I think it's about time for me to leave this floor probably.
We're a bit cashed up.
Uh, let's see. Can I use this one? Yes.
Oh, okay. I can't go there.
Love my glasses. Thank you. Have I ever tried dome keeper? Um, let me read read up a bit. Sorry.
Uh, am I giving up on learning Kubernetes?
No. I I feel like I've leed it to a decent amount. I I would need to continue in a professional setting to really get a proper feel for it. I think at the moment I like I got Kubernetes running and I was running pods on there.
So, I was pretty happy with that.
should have told you first about cube admine.
I think like that's the whole learning process.
You got to uh got to go through the frustration.
Okay. So, I don't really have any good wands here at all.
5K. What's 5K?
What's 5K?
Let's see. Can I do something a little bit better here? We've got crappy regen there.
Really bad there, too.
I could just electrify this.
That'd probably be decent.
Throw that away.
H.
What's the next chapter for me in terms of my work or something else? My life.
I do have a business idea, but I'm not uh I'm not I'm not taking any action on that right now.
Um I'm actually I am actually employed right now, but it's a bit hush hush.
This is not Terraria. This is Nita.
Oh man.
Okay, that's decent.
Thanks for the fun and good luck. Cool.
Thanks for thanks for joining in and catch you later.
Oh man, hate these sliders.
Uhoh, I'm in danger.
All right, we good.
All right, I'm not motivated to action on new business idea or just not the right time.
Um, I'm not 100% sure if the idea I have is even viable. I probably need to set up like uh like some kind of proof of concept or test. Make sure it's actually maybe maybe do a bit of like marketing first to see if or or market research first to see if it's actually going to be interesting to people.
And by the way, thank you very much for uh for joining as a member, as a distinguished member.
I appreciate it.
Uh, you now have access to a bunch of emojis that you can spam. Software or hardware biz? Uh, it'd be software.
Hardware B. I don't even know how I I would even approach that. Sounds sounds hard. Do I find learning in public to be more fun? Um, I think it's in a way more motivating, but it's much slower.
Yeah, if I was going to start a software business, it would probably be a SAS.
And uh the problem area that I would try to um tackle would be um like open- source uh stuff.
I don't want I don't want to say too much actually.
solving. You know, there's there's pain points that need to be solved, I think.
Um, so what do we got here? Kind of crap crappy perks, oil, blood.
Um, see, put this, put that crit on oil.
Okay.
Oh no, that's dangerous.
Am I staying up to watch FIFA? Oh, I'm not interested in FIFA to be honest.
You'll see a landing page test for your business idea if you said much. Yeah.
Yeah. I need to get my own uh landing page up.
See if there's any um interest in it.
NRL. No, no, I don't watch any sport to be honest.
Yes, I am in Sydney. [cough] [clears throat] Oh, electrocuted myself.
I've got a sub going for Gemini. What do you mean? I have a sub going for Gemini.
Oh my god, it's horrendous.
Oh, I've got a sub going for [music] Gemini.
What does that mean?
Oh, right here.
Oh, that's right.
I just remembered um [clears throat] the reason why I I got the subscription for Gemini was because it was included with my Google Pixel for one year, but that one year is up. So now I think I'm paying for it or that's right, I recently cancelled it because for that reason. So maybe I will have to pick up uh either a Claude or a uh Open AAI subscription. I'll probably go with Open AI because I think I like Codeex more than Anthropics models.
Yeah, this is fine.
Holy [ __ ] I'm cooked. Codeex is much better bang for your buck and better UX. Yeah, I just like how it it doesn't output as much text as anthropic models.
What do I think about vertical monitor plus horizontal monitor?
Um, I actually like one monitor.
And right now I um I actually have a very small monitor underneath my monitor.
Let me show you. Actually, I have Corsair Zenon Edge. I got this recently.
This thing.
So, that's the only the only way that I have two screens is I have it the little screen underneath. That's where my chat and my OBS is.
What's the difference between Codex and Anthropic? Codex is much better. They're the same at that scale. They are unironically the same. I have had the chance to use both.
I to be honest I can't even I can't even tell the difference between like Gemini and Claude and and whatnot.
I think Claude has the best harness where where it calls the tools well and stuff like that.
But apart from that, I just I think that anthropics are the most verbose.
And I kind of feel like they're doing that on purpose to maximize how many tokens people use and stuff. I don't know.
So that death that I just had, it ruined my win streak actually.
I think I had a four win streak.
Codex harness is better than Claude Code. Oh, interesting.
I don't think Don't you think that having multiple monitors increases your productivity? For me personally, no. My monitor is massive, so I can just have um I can put a lot on my one screen. And if I need to, I just split things in half.
I have no gold at all.
Just have a big curve monitor as your main use as a split screen. Yep. It's pretty much what I do.
Uh, these are all kind of crap.
Damn.
Okay, that perk's kind of nice.
asking again, "How are you balancing between letting AI code for you? Uh, I find it harder and harder to let it not take over the wheel." Um, I read all of the code that it puts out.
Everything.
I read everything. And um the way I approach code review, I don't know if it's I don't know if it's the same as how everyone else does it, but I check out the code on my machine and I change stuff uh until I understand what the hell it's doing. I don't just Oh my god, I'm getting mauled over here.
Man, I I do more than read the code. I read it and I have to understand what the hell it's doing. So if like it's very easy to read over a block and be like, "Oh yeah, that's obvious." But is it really obvious Of course it had to be lava.
Juniors often get into a habit of just accepting what is spits out.
Yeah, it's easy to get overwhelmed and be like, I don't understand, so but it works, so I'll just accept it. Right.
But you have to force yourself to um you have to hold a higher standard. You should be able to if someone asks you questions about the code, can you answer the questions or are you going to look embarrassed? So, I think maybe what we should try to do is think about the embarrassment if someone starts grilling you about the thing that you're pushing and you just literally can't answer any of their questions.
Maybe we need to start envisioning that embarrassment and that will provide a motivation to be like, "Oh, I actually need to understand what the hell's going on here."
Personally, once you get the fundamentals right, you start reading it like natural language. I ran into a bug the other day. Postcrist only stores microcond precision timestamp.
My code was using nanc second precision comparison.
Yeah, that sounds tough.
How does this work? I don't know, bro.
Ask Claude. And that's fine. You can ask an LLM like, "How does this code work?"
But keep in mind, it might be wrong.
And in my experience, they're often right, sometimes wrong. And I think it's the sometimes which matters.
And then it's the implications like if it works this way, what does that imply? And there's a whole trial.
Like you have to do the thinking yourself.
It's inescapable.
What are we going to do? We can't just get the robot to do everything for us, unfortunately.
Uh, this looks okay.
Okay, that's bad.
I guess I can just like spam a bunch of [ __ ] with this.
What else am I going to do really?
And then it copies that code into the project.
How does this work? What counts as fundamentals?
Yeah, fundamentals is like um I would say fundamentals is the syntax and then the standard library and then the common idioms in the language like in Rust if you're doing JSON stuff you use sirday.
It's a third party library, but that's what everyone uses.
How is this guy alive still?
Yeah.
I'm being a bit careless at the moment.
Probably need to chill syntax how it compares between languages. For example, once you verse with Go, you start picking up on how things work in different languages. Sure.
Oh, that car looks dangerous.
Thank you.
Man, not again.
Okay, well that sucks.
But I think uh it is actually uh what is it 12 uh uh 41 here which means oh man it's Monday tomorrow.
Damn.
Well, it is now that time where I will end the stream. Uh, thank you everyone who is in here.
Thank you to all the chatters. Thank you again to uh I probably won't pronounce your name correctly, but Aidu don't know how to pronounce it, but thank you for becoming a distinguished member. Very nice.
And uh we'll see. Uh I'll decide between now and and next stream whether I'm going to continue the LLM stuff. Maybe. I think I will. I'll do a bit bit more research um behind the scenes.
But um in the meantime, I'm going to go have a nice sleep.
Uh so thanks everyone for being here and chatting, contributing everything. and catch you later.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

Gremlin Arrives… While Dorothy May Takes Another Step Forward
The-moons
10K views•2026-07-23

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23

FURIOUS Raskin CORNERS DOJ over Trump DARK PAST!!!!
MeidasTouch
237K views•2026-07-23