The handoff technique is a method where AI assistants transfer context to a new conversation session by creating a summary document that captures the current state, decisions made, and remaining tasks, preventing the context window from becoming overloaded and causing the AI to hallucinate or lose track of the conversation. This technique is particularly useful when the conversation exceeds the 'smart zone' (approximately 100,000 tokens) where AI begins to hallucinate, even though the context window may support up to 1 million tokens. The handoff document should include the purpose, current state, decisions made, and remaining tasks, while excluding sensitive data like passwords or API keys.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
A técnica que engenheiros de IA usam pra ela nunca esquecer o que você pediu
Added:Come here, let me tell you something.
Your ya didn't become stupid. Do you know what happened? The context window in which you were chatting with her simply became full, overflowing, crammed shut. So, I believe you will identify with the following situation. You're there having a normal conversation with GPT, Cloud, Open Code, any of those AIs on the market, and suddenly it starts hallucinating. That's it, simply hallucinating. She stops making sense by the umpteenth message; she provides a completely controversial, twisted, and erroneous context. And that will ultimately lead you to make mistakes in your output if you take her output as absolute truth. So in this video I want to show you, teach you how to avoid this kind of thing, right, production team? The production team said that's what we need to teach here.
Exactly. So, in this video, I'm going to guide you, teach you how to avoid overloading this area and ensure that AI, Cloud Code, Cordex, and many others can deliver what you want them to deliver without cluttering the context window. Come with me. Hey everyone, Danilo Fernandes here, and today I'm going to teach you an invaluable technique to prevent the context window from becoming cluttered, ensuring that the AI delivers the results you expect. The technique is called a handoff, and in plain English it's simply a passing of the baton.
It works in any country. Hey, stay until the very end of the video, because right at the end I'm going to show you the pro version, amazing for you to just grab and insert into your LLMs, into your harnesses. Want to come with me? So, just to clarify things, today's plan is very simple.
First, we need to understand the problem, because understanding the problem makes it easier to solve. It's easier that way, right? To solve a problem that you know exists, or at least to begin the process of solving it. Then I 'll show you the solution that you on the other side probably already use and its limitations. And that brings us to one of our first questions: why does AI bore us? Well, all the work, all the dialogue, all the exchange of information takes place within a contextual window, which is the workspace. So you think about a work desk, you arrive at the office at the beginning of your day, the desk is clean, nothing, right? Just that glass of water over there, great, the PC, beautiful.
As the day goes on, you keep putting away papers. Then you put in a piece of paper from one person, a document from another, a request, some little thing, and so on. There will come a time when this context, this scenario, which in this case is your table, will become crowded. By the end of the day, you'll need to find one of those papers. If you managed to find the paper you put down last, OK, right? It's easier for you to find. But now, if you're going to retrieve a piece of paper you put down around noon, you're going to keep searching, searching, searching, and that's going to require cognitive effort on your part. And then maybe you won't find that paper, or it might take, I don't know, about 5 minutes. So AI is more or less the same thing, you know? She will have to make the cognitive effort to find, amidst so much context and so much mess, everything that you want her to find. Hey, I'm going to introduce you to a guy named Lucas. This one here is the favorite. This is our beloved, dear Lucas. Oh my beloved, no.
Where? So, this is Lucas. And here, just to give some context to where the conversation is taking place. So, every conversation with I happens in a space with a strange name that I mentioned, which is the context window. That's where working memory comes in. She is not eternal. What can Ia see now? The table where he works. So here's Iá, she has my attention. Beauty? Let's move on. So you have a clean desk at the start of the day and Lucas finds any piece of paper right away. The tension of the AI is cutting quickly, Tramontina. Beauty? That's where the inputs begin. Then you write down everything he types; everything he produces becomes an item on the desk. Each message takes up space, and obviously tokens, right? Because that is the unit of measurement for LLM consumption. They're basically chunks of words. And there you go, that's the answer.
The answer won't disappear, obviously, because it will also be added to the pile.
So, the conversation is like stacking things up on both sides. And then you realize what kind of place it is, right? Oh, Lucas's already pissed, my friend. Beauty. She answered. "Hey," she replied, "because you're like this, man."
Calm down, buddy. Awesome. Files in the stack: PDF, spreadsheet, brief. Each attachment means more paper on the table, and the token count goes up. There's Lucas stacking things up here. I'll send the spreadsheet. Now he's smiling, isn't he? Then you realize that the context is expanding. This is what AI is paying attention to. You're kind of covering it up with all this paperwork. Thanks, little bits of words. The day goes by, more questions, more answers, sticky notes, coffee, the table fills up, attention begins to wander. Oh my God! She's looking at everything. Diluted tension. The more things piled on top, the more diluted the attention becomes. So, AI will need to look at everything at the same time. Hey guys, quick message. On August 4th, at 7 PM, we'll be doing a live stream about vitalist access at C Startup. We're going to release all of our training programs, mentorships, tools, and our entire platform.
You can access it by paying once and then have lifetime access to all these updates. We'll be announcing the special launch price and all the details on August 4th at 7 PM. I'll leave a link here for you to sign up, learn more, and get first-hand access to information about our lifetime anniversary at the startup. See you there.
And then at the end of the day, where's that sheet of paper, man? So, the context window is this table where Lucas is, you can see he's really pissed off, right? The guy got really pissed off. So I brought this here for you to try and understand.
Oh, another point you can see here is that in the corner there are tokens on the table, which would be our context, right? They're already on a roll, right? So, playfully, that's what a context window is, but then there's what the big names in the industry are doing today to achieve assertiveness, something like a handoff, a passing of the baton. And what would these big names be, who would they be, the truth is they are large profiles on Reddit discussion forums. What is being discussed? It's mentioned that although the Cloud offers you, me, a context window of 1 million tokens, these users conducted a completely empirical study, okay? This needs to be very clear; it's not something that's actually proven, but according to their knowledge, around a hundred- something thousand tokens, it starts to enter what we call the dump zone, which is the "dumb zone," right? She starts to hallucinate.
Not to work, you see? So this could happen, it doesn't mean it will happen. So, this is a warning sign that we need to be more alert so we can understand if that's really the case, if that's actually happening. Because if other people in communities have started reporting this, it most likely exists.
So, even though there's a window of 1 million tokens, what's actually being used smartly, right? That would be the smart zone, the intelligent zone, is around 100,000-something tokens. That's what users are saying in posts about X and Red. And then you're going to ask yourself, okay, what can I do when the cloud, when LLM actually starts to enter the dumb zone, the stupid zone. Well, you most likely already use these resources, which is Barracct in the cloud. Yes, that's the exact command, barracct, so that you compress, so that you compact the tokens, the context window in which you're working at that moment. So the `barra compact` command will basically compress, compact the entire conversation, your files, your dialogues, your interactions, really compress it and throw you right back to the beginning of the conversation. That would be, in other words, for our smart zone. So, if you're even slightly more knowledgeable about neuroscience, neuroplasticity, neurology, everything related to the brain, you'll understand that we all function in the same way. Notice that the creation begins to truly resemble its creator. She's not exactly the same, but she's starting to look alike, in fact. The compact is nothing more than a little nap, you know?
So it is. That's basically it. When you sleep, you basically give yourself a compact bar. Then the brain will begin a cleaning process. He will begin a cleanup and will leave only the most useful information that is actually relevant. Sometimes it contains useless information, like memes, intrusive thoughts and things like that, but the idea is that it only stores what matters. And if the brain understands that the meme is important, the inclusive thoughts, then everything is fine, there's no problem at all.
So what happens the next day?
Exactly. You wake up to a clean desk and a peaceful mind. That means you had a good night's sleep, then you wake up relaxed, feeling good, go have your coffee, and so on, okay? The table, that is, a clear mind. The same thing happens here with our dear Claud and the other LLMs too, okay? But that's where compatibility limitations come in. which are basically two. I'll list two for you here. First of all, it's still the same conversation. So you compress it, then you put in another compress, then you take another compress. So you end up with summary after summary, and some information will get lost along the way. You can be sure of that. This information can be extremely relevant. If, in the middle of your conversation, in the middle of your interaction, you had to create, do a parallel task, he's going to complicate things, he's going to squeeze everything out of you, man. It's like having a small backpack that only fits shoes, and you start stuffing clean underwear in there. This is going to end badly, man. It won't be good. And that's where the fact that neither the winner nor the loser will win or lose comes in, because with compact bars on top of compact bars, everyone is going to lose.
And here I have prepared some material for you with great care. I want you to take a look. He has, you know, the beginning that we've already talked about, that we've already commented on. Full capacity isn't a useful focus, but it does have that compact and handoff aspect here. I'm going to compact it or I'm going to pass the baton. Here it is clearly explained when you will use each one. Just so you understand, we'll use the compact method when we need to summarize the entire session and still work on the same problem.
You've created a brand, let's go! You created a branch to implement a button switching feature. So there you have it, that branch is about this, it's for this purpose. So, why don't we just put a compact disc in there, right? But it's usually used for longer-term deals, you know?
So we're going to have a long conversation here, a summary, and it's the same mission, the same journey. Good for long-term debugging, ongoing investigation, and working on a single project. That's where handof comes in; it only extracts the context necessary for another mission. The original conversation remains pure, the new one is born focused.
So, the right slice, hand of. MD, because in addition to being a skill, it's also a DM file. I have been using my designs and it has been initial. So, a new mission. Good for side projects, prototypes, second opinions, and agent switching. So, let me see if there's anything here related to the agent switch that I can show you. Among the different aspects here, when the flow becomes powerful, these are the things related to advanced patterns so that you optimize your use with AI and don't lose context, don't get lost in the middle of projects, don't clutter the context window and end up paying too much for it. So the flow will become powerful in this back-and- forth movement from pattern one. The main conversation delegates a heavy exploration. The prototype only provides the lessons learned that change the decision. So, planning, open-ended questions, then there's the prototype, then the handoff back, learning without carrying the burden, among other things, I do n't know about your operation, but here it's common, the business starts with the cloud, then I go to the codex. So sometimes when I need Gemini documentation, I use Gemini.
Then there's Grock Grock 4.5, then Open Code, Kim K, and so there are several LLMs working there in different ways because the mode, the method of reasoning of each one of them is extremely different. Extremely, in quotes, but it's incredibly different. So, context window, processing power, each of them is good at something. So, nothing beats a good old handoff.m file with the proper skills. So, let me just go back here, okay? Here are four golden rules for creating a truly effective transition prompt, right? A prompt that says something like: "Dude, this is what I want.
Next time I start the conversation, I want you to start from this point." And then we have the transition prompt.
Very simple. This material will be in the description. Write a handoff document summarizing the conversation.
And here you can put the word " purpose," which is in brackets. Here you state the purpose. It includes where we left off, the decisions already made, and what remains to be done. Reference. Do not include passwords or sensitive data, obviously.
Then you come here and copy it. Don't just ask for a summary; say what the mission is for, and that the next conversation will use the document, understand? So there you have it, and that's where we get into the golden rules. I'm stating the purpose: the bridge shouldn't be duplicated. This is important because if the information is already in a file, my friend, if it's already in a file, why would I ask him to create another one? So all you have to do is point the way and not copy the intuition. Treat it as disposable. Handoff is used for shift changes. It is not permanent documentation that it can be included in the project.
Remove sensitive data from the KIS API. If you're watching this video and don't know this, it's because you're not part of our Aentic training program. Yes, information and people who build things, because inside we'll have a complete GitHub course teaching you the basics, holding your hand and saying : "My son, don't put that in your repository, okay? I already told you, don't put it there." So, this is the type of class we'll be having there, and we're looking forward to seeing you. I'm not a GitHub instructor, but Rafael will be able to guide you in the best way so that you understand that removing sensitive data is a GitHub standard, okay? So when the flow gets powerful, as we just saw, here are direct links to where they're going to send it.
Here you'll find the repository with the young man's skills, and here you'll find the specific Hand of skill, okay? OK. You clicked, and you'll be automatically transported there. So, how do I do it? How do I do it? Danilo, help me, Danilo.
You're simply going to copy this link and paste this one specifically, look, productivity handoff here.
Specifically this one here, which is already properly included in our material. It's going to press Ctrl+C, Ctrl+ V, and say something like: "Cloid, figure it out, install this skill safely." And that's where an important point comes in, which I also talk about in the Identic Builder training.
It's important to understand that, no matter how much I, Mateus, or any other teacher or instructor at No Code are recommending a resource, a skill, or a repository, it's imperative that you, as a user, don't trust us 100%. Go ahead, ask Claud to run an audit, run a complete audit, see if there are any loopholes, see if this puts me at risk in any way, you understand? Yes, that's the kind of mindset we create within the Builder training program with our students, our favorites, our no-coders. Beauty? Going back to the screen, going back to the screen here, okay? I can't read English. Oh my god.
You use a Portuguese spelling. Put a portugation. To summarize the current conversation. This is the description. A transfer document to another. So that we can continue.
Argument tip. What will it be used for in the next session? Disable invocation of the true model. So, prepare a document. You can see that it's quite simple, it's an extremely simple, extremely straightforward skill. Yeah, I already did the audit, okay? I've already audited this, so it's pretty easy to understand. Returning to our screen, you can click this link and you will be taken to this repository. If you click on this one, you'll be taken to Mat Pocock's directory. with several other skills, with several other resources, okay? OK. So here you have productivity skills, engineering skills, and everything else. I recommend you take a look, take a look over there. And here it manages the context like an identicilder. If you click, you will be taken to the link for our training and you will secure your spot to learn this and other content.
From basic to advanced. I, Danilo Fernandes, will be staying here.
I hope you found this content extremely relevant. Were you already familiar with the Handolf skill? Were you already familiar with the Holf archive? Because around here I use it a lot, you know? See you in the next video, in the next episode.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

One Must Imagine Sisyphus Happy
vlogbrothers
61K views•2026-07-21

Future of Taylor Farms
maighstirtarot5385
11K views•2026-07-21

The Downfall of OnePlus!
techwiser
65K views•2026-07-21

My Friend Locked Up The Engine On His K-Swapped Bug...
boostedboiz
128K views•2026-07-21