Mark Kashef masterfully distills the fragmented world of local LLMs into a lucid, five-layer architectural framework that empowers users to build private, high-performance systems. This tutorial is a rare bridge between technical complexity and practical sovereignty for the modern digital citizen.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Master 90% of Local AI in 45 Minutes (for normal people)
Added:Over the past few months, we've had multiple AI scares. Last week, Anthropic >> pulled access to its newly released AI models, >> Fable 5 and Mythos 5 models. Neither version is available.
>> Overnight, we lost access not just to one, but to two Frontier models. It was too dangerous for everyone to have, >> too powerful to be released to the public.
>> Now, we've got them back. But those few weeks set a precedent. It was a reminder that at any moment, whatever you do day-to-day, however you use AI, however you leverage it, could all change in the snap of a finger. Open- source or open weight models by their nature are not on par with the bleeding edge AI. But for many day-to-day use cases, they're becoming more than good enough for the work that you do and spend tons of money on closed source AI. And the best part is that they'll keep getting better. So more than ever, it's important to have the mental insurance policy of knowing how to use them, how to enable them to their fullest potential, and whether or not it might make sense to invest in hardware. So the goal of this miniourse is to break down a lot of the hairy concepts that might have held you back from wanting to explore your open-source journey in the past. There's a lot of intimidating terms and jargon that probably pushed you away, and I want to demystify as much as possible. And throughout this video, I'm going to show you a bit of a twist. I'm going to walk you through how you can use closed source models like Claude and Codeex [music] to help you set up, build, and even maintain and bug fix anything that goes wrong in your open- source setup.
So, instead of making you pick sides, I'm going to show you how to get the best of both worlds. If you watch us till the very end, then you'll be at peace because you'll be nimble, informed, and a lot more comfortable with most of the technical aspects of open source setups. Let's jump in. All right, I'm splitting this video into three core sections. The first thing we're going to do is go over the key concepts and then I'll show you how to pick the right model for the right use case at the right time and use things like tail scale or llama and others. And last but not least, I'll give you the path of least resistance to create your own open-source command center.
All right, so take a look at this. This is my personal AI command center running only on local models hosted on this Mac Mini right here. Now, for your purposes, depending on what models you want to run or what your current specs are of your hardware of your laptop or your desktop computer, you might not need one. But if you want a separate environment to isolate and run these kinds of experiments until you get the hang of it, something like a mini can be helpful. But the bottom line is that you have this dashboard with a series of functionalities that you can take off the shelf that you would have otherwise had to set up manually one by one just 24 months ago. So you can do things like having an AI chat, running and setting up Hermes agent out of the box, having document Q&A, which is basically rag, which used to be a huge headache to set up, and you have a series of other functionalities like even creating workflow automations using a local version of any. So if you click through, let's say AI chat, this will open a brand new tab using what's called open web UI. And the whole purpose of this is to give you the look and feel of chatbt or claude where you have this environment where you can drag and drop a file as you please. You can do speech to text and most importantly you can send a query and that query will run on your Mac mini not on this actual laptop.
So, if I send a question like, "What is the best type of AI consultancy for me to start in 2026?" And I send this over.
When this runs, if we go back to the dashboard, you'll see we'll have a little bit of a spike in the CPU of my Mac Mini. And this is how much memory is being used. This is the context window, and this is the primary model that we're using. And if we go back, this will obviously take longer than usual than a normal chat you'd have on the internet, but it will still give us back the result. And it will look something like this where you have the thinking patterns you can go through the response and if you want to take this on the go and when I mean on the go even on your phone if you just have tail scale set up or something like tail scale then you'll be able to interact with your dashboard even something like Hermes agent just by clicking through and it will take you to your Hermes console session you can go to chat and you can have a full conversation with this. So the beauty is not only this private mesh network that you can take anywhere, customize it to your own specific desires. And if you want to do things like image generation, you can click through here and after you set up what's called comfy UI, you can create a whole open-source pipeline to generate images like this, which aren't the highest quality like GBT2 image or Gemini 3.1, but still good enough for the majority of use cases. And on top of that, you have an extensions tab here where it's basically a small app store where you can add more open- source technology to help you with orchestration, vector databases. You have things like integrations which shows you a full map of your entire system, a full cross-section of it. And then you can see whatever models you have running on your mini. So I have a Quen visual model. I have this Quen 14B model. And I have Mistrill for things like email triage. Now the point of me showing you this is not to dangle something fancy in front of you. is just to show you where we're going to end up through the remainder of this video.
Now, before we dive deeper into any concepts, it's important to clear the air on this apparent war that's happening online. You have team closed source that thinks that open source is not only useless but way behind and can't be used for day-to-day work or at least effectively. Then you have team open source which make jokes at people only using close source saying your time is coming to an end. You will be stuck and you'll have no idea how to do your day-to-day work. As usual, you'll find me right in the middle where I think both sides are wrong. I think open source is essential to have as that backup redundancy in case all of these costs explode and you have all of these operating systems that don't need that much firepower to run day-to-day. At the same time, myself and my team use closed source AI heavily. From building custom development projects to teaching enterprises and other businesses how to use AI, I have to know how to use both worlds and both worlds have their pros and cons. So, it is a life skill in my opinion in today's age to learn how to use both effectively. Now, if you're still not convinced, let me give you four more reasons why it's worth considering running AI on a local machine. First and foremost, you have what's called vendor risk. And this is what happened with things like Fable and GPT 5.6 and will likely happen with subsequent models where some models might never see the hands and the terminals of the general public. You might have forgotten, but when we lost Fable initially and then we got it back, we didn't get the full version that we originally had back, there was noticeable differences in quality, in the length of time it could run and not make mistakes, its level of ingenuity, etc. And this is a model that is the light version of Mythos that we have yet to still see or use. Consider the bill because the subsidies of all these plans is coming to an end by nature of pure economics. If you have venture capital that is subsidizing these companies like Anthropic and like OpenAI to be able to offer $100, $200 plans that have been shown to be equivalent to $6 to $8,000 of usage on Claude and $13 to $15,000 of usage on OpenAI every single month. It doesn't take a mathematician to realize that at some point if these companies want to go public and they want their public offerings to go well, they're going to have to make a lot of money very quickly. Companies like Uber and Airbnb were loss leaders and didn't make profit for many years to get people into their ecosystem and addicted in using their products. This might be the exact same story with closed source models. So if you have an entire company who are addicted to using co-work and that is their source of truth and all their transactions run on there, all their automations run on the cloud on that platform. If the bill were to 5x, they might change over the course of a year but still they'd have to pay for it for a while because they would be locked in.
Now third is privacy and the first use case of this is obvious which is not having your personal information being trained on. So a good example is recently I built this whole health operating system and I processed all of my DNA. I had it analyzed by local models that weren't sending that information anywhere else. Once I had that synthesis and analysis then I could use closed source models to help me fine-tune and build this operating system. And the last reason is something that people either don't know about or don't know about well enough which are ambient agents. And the whole point of ambient agents is that they can constantly scour your operating system, look for security flaws, double check on the health of apps, do constant rag on any files on that system. Basically, have a 24/7 running team of agents that might not be doing the most magnificent bleeding edge activities, but are doing dirty work that you shouldn't be spending API cost to do. So instead of running open claw on even the cheapest closed source models and spending thousands a month, you would spend tens of dollars or maybe the early hundreds of dollars a month on the equivalent electricity to keep this system running.
Now some of you might have already used things like Olama in the past and you downloaded these models like Llama 3.1, Llama 2, Mistl, and you were completely underwhelmed. And you were right to be underwhelmed because 2 years ago, even 12 months ago, a lot of these models were toys. But with today's models like Kimmy K3, not only do you have better models, better systems, but you also have things called harnesses that you can attach local models to to make them work that much better. And like I said before, you can use a $20 all the way to a $200 plan on Codeex or Cloud Code to help you build the entire infrastructure. This command center that I showed you at the beginning of the video, this entire thing was set up by taking someone's GitHub repo called ODS, feeding it to Cloud Code, having it learn and fan out a bunch of agents to look through all the code, then help me set up the perfect environment, tailor it to me, find the best model for my specs, and take care of all the plumbing. So, the goal here is we can use the easy button to creating our local open-source system. Now, this is the most important point of the whole video. And once you understand this ecosystem diagram, then everything else will fall into place. So, step one, if we start from the back all the way to the front, you have your hardware layer.
And this could be your Windows laptop, your Mac, your Mac Mini, your Mac Studio, whatever it is that you own, this decides pretty much what models you could run, at what capacity, and at what latency, or what we call tokens per second. Once we understand what we're dealing with, then we can understand which models. And when we refer to local models, some people might say open source models, local source models, or open weight models. Open weight just means that we have the entire file of weights that represent how these models tick. So with things like Anthropics Claude and OpenAI's GPT models, you don't actually understand how they work.
They don't have the weights publicly available. But with these Chinese models, GMA models from Google, you can physically see these weights. So you can see exactly how it deres the answer it comes up with. Once you know what models you can use, you then have your inference engine layer. And this is basically involving something called Llama CVP usually or VLM. Llama CVP is basically the bridge through which you have local models take the input information or the prompt from you run it through this engine and come up with the output. And VLM is just an alternative that is used more at a team scale. Once you have your engine established and that's all taken care of, then you get to what's called the gateway portion. And the gateway is typically something called light LLM.
And this allows you to create what are called APIs that allow you to connect any service to these local models running on your computer. So you can think of it as a plug adapter that allows all of your applications to talk to whatever model you want. And once you have the gateway, you have the final layer on top which allows you to interact with these models through platforms like open web UI where you can have a back and forth just like you're using chatbt or claude. And on top of that, if you want to make these models more effective, then you hook them up to what are called harnesses. And harness is basically allowing you to have this open source model as a brain in a jar.
Harness allows you to attach limbs to this brain. So it can do things like write, read, manipulate files, use things like bash to control and move files around your computer, etc. And this harness is the very reason why people love clawed code and codecs. They think that all of the orchestration is the model itself, but it's actually the symbiosis of the model with all of its limbs that it uses to actually manipulate its environment. And outside of that, if you want to be able to use the model with something like Cloud Code, you can always hook up Open Router through API and use Cloud Code's harness and borrow it with open source models.
Or you can do this locally as well. And by the way, if you enjoy the way I teach and break down some of these more sticky concepts, this is what I do day in and day out in my early AI adopters community. I go through all kinds of concepts from how to use cloud code to codeex toentic harnesses to even how to monetize AI responsibly in this day and age. So, if you're looking for a one-stop shop without the hype, without the obsession of things that don't actually matter and move the needle, then check out the first link down below and maybe I'll see you inside. All right, back to the video. So let's tackle all of these layers and what we can do is bundle the first and second layer together. So we can use something like a cloud code, a codeex or even something like a Kimmy 3 to look at your current hardware and understand what is your RAM, what is your memory, what does that mean in plain English, what is your VRAM and then understand based on this what are the best models that you can run where they'll actually practically run and you won't wait five business days for an answer to hello.
So if we hop into a terminal, we can send over a very basic prompt like this.
And I'm using Opus 4.8. And we can say, can we use open router API to pull the latest open source models in the past 6 months as of today's date. And then we can say categorize it into texton generation models versus multimodal.
There are even some models that aren't that efficient, but they let you run completely for free. Not the most performative, but they can actually work. Now, before I keep going, I want to zero in on what Open Router is, just in case you don't know. It's a platform that pretty much allows you to talk to any model, whether closed or open source. You would basically have this credit wallet, and you can pull from that credit wallet by using that API.
So, you can access all kinds of models from one main bridge. And the core thing we're using it for is this explore models feature where it has rankings, an entire leaderboard, and they have a public API that you can tap into. So in this case, all we're doing is pulling the comprehensive list from that source as our proxy to avoid us to have to do something like web fetch and researching where sometimes it's on and off on how comprehensive it is. So once we get that full list, it will dump it as a JSON file. You can see here all of the models and in this case we care about text only versus multimodal because if you want to be able to generate images and understand images, create videos and understand video and create audio and vice versa, you want to know exactly what is the right model for what you're looking for. And one of the benefits they have on top of having an entire database of the latest and greatest models is this rankings tab where you have all the top models, a whole leaderboard, top models by task. So if you want to be able to know which models can consume and create images are multimodal. This would be a very easy way to understand that which are great at debugging planning and obviously you can see anthropic owns a lot of the leaderboard here. But you can then filter on open source models see which ones would fit the best for your specific tasks. But with that we can have cloud code figure out how to use the open public API and process the information. So that's what it does here. It pulls the information. You can see here all of the textbased models and then it categorizes it in these tables.
So these are all of the textonly generation models in the past 6 months.
And then we have the multimodal ones, the ones that can sometimes understand and create things like image, video and audio. So this is the full list and this is the key thing here. Now we have the list and we could say something like this. I want you to go through my current operating system specs and tell me what models I'm eligible to run if I wanted a model that's really good at generating proper and quality text, but also can understand image and potentially video and create audio. If I need multiple models to accomplish this, then let me know. But I want to make sure that it doesn't run very slow on this computer. Now, with this seemingly vague prompt, you can send this over and it will be able to access the specs of your operating system, then compare that to what these models need, and give you the right short list to get started.
Now, if you're curious what I'm running on, this is an Apple M5 Max with over 128 GB of memory, 18 cores, a GPU, and overall, it's a beast, but it's very expensive. You will not need this necessarily for whatever you want to run day-to-day. I wanted to make sure that I could invest in all the hardware I could before all the prices inevitably rose a couple months ago. So, if we scroll down, it will break down exactly depending on your laptop and your OS what makes sense for you. So, it breaks down the chip, the unified memory. And really important of this memory, how much of that memory can be used towards models. Sometimes people will read the box of their laptop and say, "Oh, there's 128 or there's 50. That means I can use 50." you can't necessarily use that because you have other apps that are running on your computer and you need to take that into account. Then you have storage and then it breaks down exactly what models might make sense for me. So you'll see here I can run the Quen 3.5 3.6 GMA 34 31 bit. If you don't know what the word bit means, we will get to that in a bit. Pun intended. And then you have the Nvidia Neotron. And then it tells you which ones are way too big to run. So although this is a beast, it's not beast enough to run something like a Kimmy K2. You would need what's called a quantized version, which we'll talk about later, but it's basically a compressed version of that model that is meant to run, but maybe not run at the fullest capacity. Now, in this case, it couldn't see which models have the best text to speech. So, one thing you could do is ask it to spin up or fan out a series of agents or sub aents to look for the best models to allow you to create and do text to speech. And one additional thing here is it tells you exactly how you can run this. So it says you can run both simultaneously. And the easiest way to serve these locally on Apple Silicon is LM Studio or O Lama.
And you can see right here Llama CBP backend. So the wrapper around this specific engine. And then there's something called MLX that's optimized for Apple products. Now if you want to be able to download these without adding additional software, then you can literally say, can we download this specific quen model that you recommended? It will check what's called hugging face. And hugging face is basically the equivalent of an app store for models. So whether they are brand new models, whether they're quantized, meaning compressed versions of existing models, or fine-tuned models, you'll find them all here. So open router can serve it, but you can't necessarily download it. Hugging face is this entire repository where you can download the models. These are models released by vendors themselves. Let's say like Kimmy, Quinn, GLM, as well as open- source models that have been distilled.
Meaning maybe someone took a version of Kimmy and not only compressed it, but made it really great at understanding images for some reason. You'll find all of these different permutations of models here and it was smart enough to know to have to investigate and see is it downloadable from this database. If we scroll to the bottom, you'll see yes, it's downloadable, public, and there's a version tailored for your Mac. Here's what I found. So, in this case, it will give us some options. I will go with whatever is recommended. You just want to double check that you have more than enough room to support this because in this case it is 70 GB and it seems like it has both image and video reasonably fast and it looks like we can handle it.
So I will click on enter. This will start the downloading process. Now in my case I already have O Lama and Brew installed on my computer but if it's missing from yours it should be able to understand that it's not there and install it as a prerequisite. So then it could set up the download process and you'll see this will take a little bit of time. It will install it, optimize it, and then it will be ready to use on this laptop. It not only pulls and downloads the model, but the best part of running this terminal with a closed source model as smart as Claude is it realizes that I need to update my version of Python to run it. You can see right here the 65 GB download was perfect, but the serving library needed two fixes. And instead of me having to figure out what that means, it would go and update what's called your virtual environment and then the model architecture to make sure it runs as well as possible on my specific computer. And if we scroll to the very bottom, it tells you that it's fully operational. It tells you the number of tokens per second. So this model will work at 60 tokens per second. This is exactly how much memory it will take of my 128. Typically, you don't want to exceed more than 50%. you'll start running into problems, especially if you're running other things in parallel.
And if we scroll to the bottom, I had one issue. This one issue was I wanted to be able to interface with it using something like LM Studio or Olama. Now, if you remember, because both Olama and LM Studio are wrappers around Llama CBP, Claude Code was able to stitch together the connection to this model. So, now we can load it and we can actually have a conversation with it. Now, hopefully the video doesn't collapse, but if I say, "Hi, how are you?" We should get, like it said, 60 tokens per second. We can already see that it's thinking. We could take a look at its thinking traces and we should get hopefully a very simple response fairly quickly. And I'm running this while I'm running this video in 4K.
So, it's doing a decent job. Now, before we keep going, I'm going to eject the model so I can make sure we have as much memory as possible. We'll close this.
We'll close. Now, the same way we did this locally, we can also do this on my Mac Mini. I can connect to it through screen sharing through what's called SSH secure shell and I can run that exact same prompt over there where I can say as of today's date find the best performing models for text to speech on the specific hardware and now you know that the whole pattern it will scroll down it will find the best matches it'll create recommendations I can tell it to install it on my behalf and then once it's done it looks like it's pretty much close to that I can always tell it go and update our dashboard to make sure that I can see this under the models tab if it doesn't happen automatically. So, we've already tackled three of the five layers and before we move on to the gateway layer, I just want to close the loop on what the difference is between Llama CBP versus VLM. So, if you think about it mentally, when we are working with a localbased model on our personal laptops or hardware, you're basically cooking for a party of one. When you're cooking for a party of one, you don't have to worry as much about scaling that inference across the board. But if you are serving the equivalent of an industrial kitchen, you need a way that this can scale and multiple people can hit the same local models at the same time. And that's where something like VLM comes in. It basically allows you to do what we did locally, but apply it at scale. Now, another way you can think about this is through taking the analogy of a car for yourself. You might be more than fine with driving a beater that will take you from point A to point B, even if it's a little bit slower. it has some weird sounds on the road. As long as you get the output you're expecting.
But as soon as you're forced to drive that same car on a highway with a speed limit of 300 mph, you're going to have problems. So the point of VLM is it allows you to run these engines at a much higher capacity so that multiple people can use it concurrently. Now let's take a look at the penultimate layer which is the gateway. And one of those gateways is called light LLM. You can think of light LLM as a switchboard.
And the whole point of it is that if you want to be able to interface with your local models through things like chat through the open web UI through things like agents that can either code in parallel asynchronously or run in the background as ambient agents or run automations through naden hosted locally all of them can go through this middleman which is the light LM layer.
And the beauty of it is that it has what is called an OpenAI compatible API which allows you to have one backend that can speak to hundreds of models. So all you have to do instead of having to download a model like you saw right now and constantly rewire it to the chat or to work with Olama or everything else, you have this middle layer that helps with the orchestration itself. So if you were using Quinn and you want to move on to using GMA, then all you'd have to do is swap this cable for this cable right here. And the whole point of this is to keep you as nimble as possible.
Now, before we come close to the last layer, which is actually setting up the final interface that hooks up to everything running on your system, we have to understand a few more key concepts. One of those concepts is the private mesh. And this very mesh is the reason why I can access this dashboard that is running infrastructure on my mini. And the way it works is it's using something called tail scale. And tails scale is basically a private encrypted network that makes it possible to have one central hub where I can tap into every single device as if they were sitting in the same room. So instead of doing things like port forwarding or hosting a server where someone could possibly compromise it and get access to those devices, you basically have them all chat to each other secretly in a way that nobody else can see. And one of the many reasons why I like it is it's not only very easy to set up, but also when it comes to pricing, and again, no affiliation here whatsoever, is you can use the personal plan completely free.
And that gives you access to unlimited user devices. So whether I connect one Mac Mini or five, it'll be the same price, and you can add up to six users, which for even small businesses, very small businesses, this is more than fine. And not only is it economical, but it's also very easy to download. You can just say tail scale download app Mac.
This will take you directly to the download page and this will give you this little app that sits on the very top of my screen. But if you click on this and you click on network devices and my devices, you'll see right here I have my iPhone, my old laptop, and the Mac Mini all hooked up together, including this specific laptop. So if I were to run an AI chat from my phone and I go on Quen 14B and I just say, "How are you?" And then we send this over.
This is what's happening behind the scenes. Like I said, we have an M2, we have an M5 laptop, we have an iPhone and a Mac Mini in this network. When we send that request from the phone, it will directly travel encrypted to the Mac Mini. The Mac Mini will actually run the inference. It will run the engine and then it will transport back the response securely back to the phone on that dashboard. So the whole point of this is it allows you to take your entire network on the go. You can integrate it with all kinds of tools and you can even share it with someone else if you want them on your private network. So now that we're more comfortable with the architecture, let's go back and look at the best models as of today. And naturally, just like the close source space, by the end of August, by the end of September, by whenever you end up watching this, the model names might have changed or the model versions, but most likely the providers will be similar. So from left to right, we have GMA, which is released by Google. It's a decent model, but sometimes it's weirdly prohibitive. it will stop you from doing very normal legal things for whatever reason. Once you get to the Chinese models, I'm going to say the quiet part out loud. They really outperform a lot of what's on the market. So whether it's Quen, GLM, Kimmy, not only are they releasing brand new versions all the time, but their token windows are now reaching that million context window limit and their costs if you host the entire model are actually going up. But to host the quantized versions, the compressed versions is increasingly cheaper. Now, obviously, DeepSeek gave us the Deepseek moment a few years ago where open- source really became a part of the conversation and Miniaax, interestingly, offers alternatives to things like 11 Labs. So, when it comes to voice cloning, it has very unique characteristics versus the other models.
If I had to pick my favorites, having used all of them now, the Quen models and miniacs give you really great agentic coding options, but also you can use GLM as a workhorse. So let's say you're using a cloud code or a codeex and you want to be able to still pay for closed source, but maybe offload admin tasks to workhorse models. You can use both where you have the close source, take care of really detailed planning, bug review, testing, and then for small tasks like go and change all the names of these files, go and search for this specific file using GP, which basically lets you search through a variety of databases, files, etc. You can start to be more intentional with what you offload to what kind of model so you can save on cost and tokens. So naturally, you can try before you buy and you can use something like open router that we talked about before to test out all these different models on a series of tasks. So you could hook it up to cloud code or codeex natively, which I'll show you in a second. And then you could test five different models on 10 different tasks and see which ones perform the best on the workflows of choice. Once you have that, then it's easy for you to decide what deserves to live on this local hardware permanently. What we can do is say something like cloud code integration open router docs. Then we send this over. You'll see the first link is cloud code integration. We'll click on that. You can read it yourself.
You can feed in the page or for me I'll just feed in the raw URL. So I'll go into a brand new session in cloud code and I'll say something like I want to be able to use open routers API natively through cloud code. read through all the documentation, spin up any sub aents if you wish to look at any details and create a full plan on exactly how we can communicate through the API maybe through switching the model to an alias that's called open source. So we'll send that over and then I will paste this link. It should be able to now start using its web fetch tool, pull the information and create a full plan. So after a few minutes it comes back with a full assessment of our options and if we scroll to the bottom it tells me exactly what I need to decide on. So what is the primary open source model that I want to use? And it tells me that one way it could implement it is to create a specific mode that's called claude OS.
If we pick cloud OS then it will run using the open source models. Otherwise if you use claude then it will use the standard models out of the box. So in this case I'll just say let's scroll down and let's go to something and say I want Kimmy K3 and then how should I handle your open router API key. Now multiple ways you could do this. You could use the keychain. You could use a private source file. Now, for simplicity, I'll just use a file. We'll do isolation. We'll send this over. And it should be able to come back in 5 10 15 minutes and hopefully have this set up. And just like that, everything is set up. And it created a shortcut command that's called DSPOS. And DSP stands for dangerously skip permissions, which is basically running Claude Codes harness in YOLO mode. So, if I go into a brand new terminal and I write DSP-OS, this will open up Claude and then you can see right here it's using moonshot Kimmy K3. If I want to be able to switch the model, all I'd need to do is write the same command but with the name of the language model right after. So, if we pop back over, it created this other shortcut command that's called DSPOSG glm. If we run this in a new terminal, doos one more there. GLM. It should ask me permission. And there we go. It's now using GLM 4.6. And if we say build me a tic-tactoe app, what's happening behind the scenes is that this is being sent to the open router API. It's interfacing with this model. You can see right here, it's entering plan mode and it's borrowing the harness of cloud code. So, you get the best of both worlds. if you want to be able to use this kind of model at a lower rate of inference, see how it works before you decide to actually download it on your hardware and using it all over the place. And one thing I wanted to show you is you can see that it's executing this task and at the very bottom here, it's running some agents and sub aents. So even though it's still not the cloud code model, because we're using the harness, it has access to all the tools that we know and love that we typically use closed source models for. Now, leading up to the final build, I'm going to go over some key concepts in quick succession of words you've either heard of before and have no idea what they mean, or you might think you know what they mean, but you want to be 100% sure. One of those concepts are the number of parameters of a certain model, which are usually in the billions. So, if we pop over to our command center and we hover over the AI chat and we go back to our model, you'll see we're running quen 3, but we're running the 14 billion parameter model.
What does that mean in plain English?
Now, for all intents and purposes, you can assume that the number of parameters is a proxy for how capable that model is. So, if we take a model like Kim K3, it has almost 2.8 trillion parameters at the full scale model. Running that at home could cost you $10,000 a month plus in pure electricity costs, let alone all the hardware you'd need to use it. So, when you compress these models, you make them less capable. But at some point, these models have become good enough that a 14B, a 27B, even a 50B can be more than sufficient for a lot of your day-to-day work. And this will keep increasing over time as closed source improves and open source improves in tandem. And one rule of thumb that you can use is you can take the total number of parameters in a model, divide it by roughly two, and that will give you an approximate amount of how much memory you need on your computer to run it freely. So if we take the 70 billion parameter model that we downloaded earlier, it needs roughly between 35 to 40 GB of RAM depending on your computer and a few other factors. But overall you can start to see that if you run 120, you would need almost 60. And if you have a very small, not as agile computer, then 14 would need anywhere between 5 to 7, which could be feasible.
Now, one thing we didn't touch on is what 4bit means. It's pretty much a proxy for quantization and quantization is a proxy for compression. So if we jump into quantization, imagine you have a very smart brain and as you compress its knowledge and its fidelity of knowledge over time, it's kind of like resolution of an image or resolution of this video. This is being streamed in 4K. It could be in 1080p. It could be in 720. Once we get to 480, it becomes really rough. And the exact same thing happens with compression. So on average you could have 16 bit, you could have 8 bit, you could have 4bit, and you can continue until you go to one and below that. And at that point, once you get to one and below, you've overcompressed the brain to the point where a lot of the intelligence, a lot of the capability, the ability to know how to use tools, even with a great harness, become severely prohibited. And if you ever see any one of these words in the wild, I don't care to go through these compression algorithms specifically, but just know that they're different versions of compression. So you have a combination of how many bits and what compression algorithm and then you come down to the number of billion parameter models. So what you should care about if you're running things locally is finding the smallest quantization of the smallest billion parameter model that still passes your tests and accomplishes whatever it is that you're looking to do in your day-to-day workflow. Now the next key terms here are GGUF and MLX.
You can think of GGUF as this universal passport. It can work with pretty much anything. If you use a GGUF model, of which you could find tons of them on Hugging Face, it is the most generic and versatile type of model. So, if you download this, nine times out of 10, it will work with Llama CPP, meaning it will work out of the box with O Lama and LM Studio. When it comes to MLX, this is basically an Apple native passport. So, the same way that GGUF works with all gates, MLX is specific to Apple. So, it is meant to work with Apple Silicon. So you can use the same model, but it will perform much better because it's optimized for that specific stack and architecture. So the rule here is simple. If you're not running on Apple, then you have no reason to look for GGUF. And of all the models you'll find on Hugging Face, you usually will find an equivalent of MLX to GGUF for every model. But for some, they just have the general version. So that's one thing to keep in mind. You can run GGUF on everything, including a Mac OS, but this one is specific to this operating system. Now, this next term trips a lot of people up and it's called KV cash.
And the way you can think of it is when you start a conversation, you're going to have a blank slate or in this case a blank desk. As you have more back and forth and you have a multi-turn conversation and potentially that multi-turn conversation is calling tools, maybe it's using a harness. Now, you have the system prompt of the agent you're using. Let's say it's Hermes agent, which has a mammoth system prompt. And then you have back and forth conversations. As you keep going, you accumulate tons of context. So this desk here on the lefth hand side keeps accumulating. So one way to optimize the use and the performance of local models is to manage your KB cache. This could mean adding or basically distilling a shorter system prompt for whatever agent you're running. Now of three ways that you can optimize this cache, one of them is to have a shorter system prompt for your agent. So exhibit A would be for our Hermes agent. When we clicked into it and I chatted with it using the 14 billion parameter model on my Mac Mini, it took around 2 minutes to respond to high and then I ran a diagnostic and this diagnostic showed that the system prompt being pushed through the model was humongous and a lot of it I didn't actually need. So as soon as I pruned it, I went down from 2 minutes for a hello to 30 seconds. So doing this kind of pruning on the system prompt is helpful. And then obviously adding smaller context meaning if you don't need to write a whole essay or a story for a specific instruction then avoid it. And lastly similar to quantizing a model you can also quantize or compress the cache. So as you have a conversation which is something that a lot of chat agents and chat bots do in the wild.
Maybe after 20 messages you summarize those 20 messages as one message and only that one message is persisted in the next parts of the conversation. And this keeps going to give your conversation as much bandwidth as possible. Now there's something that's called sparse attention that became popular with the mini max local model.
And the way it works is as follows. On the left [snorts] hand side you have your standard model where it reads every single word of every single part of the conversation at all times. On the right hand side you have sparse attention where it picks and chooses which tokens might be the most relevant when it looks back at the past. So this makes it a lot more token efficient and helps alleviate a lot of the performance issues that can come with larger models. Now before I briefly touch on harnesses, I want to mention one last term which is tokens per second. And the way you can think of this is as follows. Your memory bandwidth on your computer is like a pipe. And the width of the pipe allows it to dictate how large of a block or blocks can pass through and at what speed. The smaller the blocks or the smaller the bits in this case, the faster they can run. If they're very large blocks in a very tight pipe, then you're going to have anywhere between 5 to 10 tokens per second, which is actually very slow. That's the equivalent of getting five words over the span of maybe, let's say, a minute.
You'd almost have a paragraph. Whereas with 60 to 70 tokens per second, now you can start to have a paragraph within 10 to 15 seconds. So, in terms of the quality of life, having a more quantized model allows more agility and allows you to have a better experience on your specific hardware. And if you're curious on how speed is derived, this is the equation for it. I won't dive into it just in case you're non-technical and if you've already had enough. When it comes to harnesses, we already touched on being able to use something like cloud code and combining it with open router.
You can also find ways to pipe things like Olama and use that exact harness locally without the open router. You can use codeex's harness. You could use Geminis's and you can also use open-source harnesses like pi.dev. And the way this works is it's a very minimalistic, very malleable framework that allows you to start to build the limbs that you can hook up to your brain of choice. Now, you can set this up by copying and pasting this curl command and running it in a terminal and following all the onboarding. Or you can do what I do and give Claude code or codeex the link, tell it to read the documentation, fan out a series of sub aents, become familiar with it, then set it up for your specific operating system. Now the key trick with this is once you have it initially set up it will not be as robust as cloud codes or codeexes but it's going to be your harness. Meaning as you monitor performance keep giving feedback to your language model of choice to improve the harness because after 5 10 15 iterations it might get to the point where it's optimal. Now the main trade-off with using pi.dev versus something like a cloud code harness is if you just want to use close source models a lot of the harness engineering is done for you.
they keep updating it to be optimized to work with the right model at the right time with the right tool calls. But with pi.dev, as you keep having brand new models come out, you will incrementally need to keep tweaking it because all models will have slightly different behaviors and different patterns at which they like to call tools. Now, with all this information, you have more than enough to go to the last step, which is to set up your local command center. And there are all kinds of frameworks that you can choose. I'm just going to show you one that worked really well for me.
Now, if you search for ODS GitHub, this will take you to this Osmantic ODS. And this is a newer repo by someone named Ahmed Ozman. He's actually very gifted at this open source stuff. I'm giving him a huge shout out cuz he's done incredible work with this. And he's open sourced this entire command center that I've built on called Osmantic. Now, using the same technique that I've shown you thus far, which is taking repos and websites and documents and feeding them to your language model of choice, you can take this GitHub link, paste it into your language model of choice and ask it to one, inspect it and audit it, and then fan out some sub agents to plan how you can implement it on your specific computer. And this is exactly what I did. So I said fan out some sub agents to read through this GitHub repository to comprehensively come back with a plan on how to set this up locally on our system with our downloaded Quen models.
So this is the GitHub link. It then does some web search. It goes through the entire repository. It looks through everything we might need to make this work. It comes back with its initial findings. So in my case, we have an MLX model that we downloaded. So, we just need to make sure that it works well with the llama servers and any issues that pop up. Even if you're non-technical, you typically just have to speak to it in plain English, get to the bottom of basically what is the path of least resistance. And once you understand that, it can go and actually execute all the steps. So, I was secretly running this the entire time I was filming this video, and it took around 45 minutes with a couple queries on preferences. And once it's set up, it now should work out of the box. If I go to here and I put this local host address, it should initialize, ask me to label the platform. So I'll say marks ODS or maybe I'll just keep it MarkX ODS. There we go. We'll click on continue. I'll just say username is Mark. Continue. Let's do everything and then continue. We'll click finish. And here's your dashboard. And the core differences between what I showed you at the beginning and this is one, we don't have the logo at the top left hand side.
Two, I enabled all the different features here by going back and forth with cloud code. And this is using instead of the out of the box model that it comes out with, which is the fee model, it is using the one that we downloaded earlier. So now we have a full pulse on exactly what's happening in our system, where everything's located. We can start to install the extensions, add the integrations, and you can see we can start to really stack this up and get it to be as sophisticated as this one is here. And naturally, you can take this and combine it and Frankenstein it with a series of other frameworks to make your perfect snowflake version of your local setup.
So hopefully this entire walkthrough gave you the full TLDDR of exactly how you can go from not knowing how to use local models to using them actively and using them effectively with the path of least resistance. And I know I threw a lot of you in this miniourse. So, on top of giving you all the documentation that you need to review everything that I went through, I've also put together this local AI engineering [music] guide.
It is extremely sophisticated. It documents everything that I've read, seen, and surveyed in the past month and a half that I've been setting up my own rig. So, you'll get both of these resources along with a few other goodies completely for free down in the second link below. And as always, if you want to level up your AI game beyond just local models, but also how to use things like Cloud Code and Codeex to their fullest [music] potential and be supported by a team of 12 plus coaches, myself, an entire community, and all the other resources that you'll never see on YouTube, then make sure to check the first link down below and maybe I'll see you in my early AI adopters community.
And for the rest of you, if this gave you a good primer on how to use local models and you feel that much more confident, I would beyond appreciate a like on the video and hopefully a comment to expand the reach. I'll see you all in the next
Related Videos

TOP 15 Data compression Interview Questions and Answers 2019 Part-2 | Data compression | Wisdom jobs
wisdomjobs
281 views•2019-06-28

CTS 158: 802.11w Management Frame Protection
ClearToSend
4K views•2019-02-04

NDSS 2019 Send Hardest Problems My Way: Probabilistic Path Prioritization for Hybrid Fuzzing
NDSSSymposium
496 views•2019-04-02

How realistic is Cities: Skylines?
CityBeautiful
159K views•2019-02-14

GUIs & TUIs: Choosing a User Interface for Your Python Project | Real Python Podcast
realpython
2K views•2025-04-04

The OSI Model - Explained by Example
hnasr
225K views•2019-05-12

Cloud Computing - Introduction
elithecomputerguy
98K views•2019-10-07

From Traveler's Dilemma to Dynamic Routing | Demystifying Networking
IITBombayJuly
5K views•2019-08-04
Trending

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Steam and Xbox Just Dropped The Hammer On PlayStation
OhNoItsAlexx
9K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23

SuperBike Factory Has Gone... What's Next for the Motorcycle Industry?
thatbikersimon
11K views•2026-07-22