Andrej Karpathy's Auto Research demonstrates that AI agents can autonomously run continuous experiments to achieve goals without human intervention, using a looping system where Claude repeatedly tests and improves work until meeting specified criteria. This approach overcomes the limitation of single-prompt interactions by allowing AI to iteratively refine outputs based on custom evaluation metrics, such as optimizing website speed, improving email content through multiple scoring cartridges, or enhancing race performance through systematic variable adjustments. The key principle is providing AI with clear objectives, evaluation criteria, and boundaries, then allowing it to work independently until the goal is achieved.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Stop prompting Claude. Use Andrej Karpathy's Loops Instead...
Added:So, Andre Karpathy now believes that prompting Claude is outdated. Instead, imagine telling your Claude your business goals, then walking away and letting it test and improve its own work until it achieves the goal that you gave it. Well, that's what Andre Karpathy created with Auto Research, a looping agent that autonomously ran 650 experiments across 2 days while Karpathy watched Singles Inferno. Well, today I'll show you how you can build your own Karpathy-style loop inside your business so we can all stop prompting it like a pleb and get 10 times the power out of our Claude.
So, let's start with the problem with prompts.
>> To get the most out of the tools that have become available now, you have to remove yourself as the as the bottleneck. You can't be there to prompt the next thing. You're You need to take yourself outside.
>> Right now, most people are using their Claude like this. They open it up, they type a task into it, Claude gets the task, goes and gives it one attempt at the task, and returns it back to you being like, "Hey, I did what you said."
Sometimes the results are incredible, sometimes pretty mediocre, and sometimes we question why we even pay for Claude in the first place. But, the problem is not Claude. It's how we are using him.
We only gave him one attempt to do what we asked. Imagine we had a developer as an employee and we said, "Hey, can you speed up this website for us?" You give it to him, he makes it 10% faster, he sends it back, he says, "I'm done." Now, it's not the fastest website possible, but it did satisfy the request we made of it. This is what our Claude is currently doing with all of our prompts.
And this is where the beauty of Karpathy's loop comes into play.
Karpathy gave his Claude one clear goal, to speed up his website. But, instead of giving it just a single prompt, he looped it. Meaning, every 5 minutes, the agent would try again to make the website faster and faster and faster, building upon its past results. And over the 2 days of this loop running in the background without Karpathy having to be in the mix, his website got way faster even after Karpathy personally tried to speed it up himself, and he thought he reached the cap of how fast a website could be.
>> Yeah.
>> And I'd gotten to a certain point and I thought it was like fairly well tuned and then I let our research go for like overnight and it came back with tunings that I didn't see.
>> In fact, Andrej Karpathy thinks that it's funny that our human meat suits are in the way of agents doing better and better jobs as it is.
>> The name of the game is how can you get more agents running for longer periods of time without your involvement doing stuff on your behalf. I don't want to be like the researcher in loop like looking at results etc. like I'm I'm holding the system back.
>> This is his auto research. GitHub, we're not going to go into the technical side of this, but it's this fascinating paragraph right here that says, "One day, frontier AI research used to be done by meat computers between eating, sleeping, having fun, and synchronizing once in a while using sound waves talking back and forth, interconnecting in the ritual of a group meeting. That era is long gone. Research is now entirely the domain of autonomous swarms of AI agents running across compute cluster mega structures in the skies."
This looping feature is the start of this future. Utopian future in many people's eyes, does sound pretty dystopian based on what we've been taught so far though. The agents claim that we are now in the 10,000th generation of the code base. In any case, no one could tell if that's right or wrong as the code is now a self-modifying binary that has grown beyond human comprehension. And he says this repo or this loop feature called auto research is how it all began. This is March 2026 where he tweeted this out.
But since March 2026, Andrej Karpathy's auto research has come a long way. Now, they are actually stored inside Claude as innate features. There's {forward slash} loop where you can wake Claude up repeatedly to check on something, and there's {forward slash} goal which keeps Claude working until the clear condition is met, which is this distilled down version of Karpathy's auto research.
>> Auto research is just, yeah, here's an objective, here's a metric, here's your boundaries of what you can and cannot do, and go.
>> But the biggest difference is that your Claude will work on whatever task you give it until you tell it to stop, optimizing it over and over. So, if you're working on social media content, the social media content's going to get better and better until the human, or you with your big old meat suit, get back in there and tell it, "Okay, that's enough. It's as good as I want it to be now." Or your website will keep getting faster and faster until you jump back in the loop being like, "Hey, good job, Claude. We're done here." I'm going to show you two live examples that we're working through together of a Kapathy loop. The first one I want it to be extremely visual, so it's a race simulator. And the second one is a real-world business loop integration.
Starting with our race simulator. So, we have this little guy here racing around the track, but he hasn't learned how to drive. And there's a bunch of variables that go into how quick his lap time is going to be, such as the driver, the wheels, the turning, the weather. And we want to put Claude as a race engineer to find the fastest and cleanest lap around the track to improve his time. So, he started with a time of 60.08 seconds per lap. But, we're going to get a Kapathy loop onto him to speed him up.
I open up my Claude code, and I type forward slash goal. Then I say read Claude prompt.md, which is actually the master prompt to set up this loop. So, we are telling Claude that they are the autonomous race engineer for the Apex loop. Your job is to search for the best complete race package through a Kapathy style experiment loop. The goal, keep improving the race package until all three deterministic seeded runs finish.
The median lap is to be in the goal here below 40 seconds. Remember, it started at 60 seconds. So, we have a few more things in the prompt here that are the rules to set it up properly. But, coming back to this prompt, then run the Apex loop race package experiment exactly as written. Follow its four controlled experiment phases, changing only the driver in driver.js. Remember, we it can't change how we're actually scoring this. This is the most important thing of a Kapathy loop is that the agent only has access to the file or the thing that you want to improve. And do this until you reach a time under 40 seconds. So, we're going to send that off to our Claude code and start our Kapathy style loop. So, our Kapathy loop has finished and you can see where it started. It studied the physics engine, it then tracked the model and geometry. I had a complete model of the physics and then he went into his first experiment, which is the baseline. Run all three runs with the medians. He's got his stats and data here and unfortunately that was really bad. I'm going to show you the results in just a second. He went on to experiment two telling us what he's optimizing for and what he's changing every time. Remember, there was no human in the loop. I did not touch it after I set this loop off. He then went on to experiment three, which was uh improving another variable that we had and experiment four as well. So, when we scroll down and we see these in a table, you can see the baseline was 60 seconds.
His first experiment actually made it worse with more collisions. Then, he had a predictive safety package installed.
Remember, no human in the loop, no meat suit touching the keyboard at all. And he took off 23 seconds of the lap time and then his fourth experiment was even better again getting it under the goal of 40 seconds, which is where the loop stopped. And you can see this is the new and improved driver package thanks to our Claude code loop displayed visually here on the Apex loop racetrack, which is really fun in theory. But, I also like putting these into real world practices for us so we can actually use them for our business and use them for our personal life and have these Kapathy loops set up at all times running hundreds of sub agents working towards our goals for us while we go watch Singles Inferno. Okay, so for a real world example, this is Claude's official diagram of a goal-based loop. He came out with this on Twitter. Claude will work on the task and once it's done its work, it's going to try to stop. But, it only gets to stop the task once it reaches the evaluating model's condition that you've given it. In the last experiment we just did together, it was a 40-second lap time. In this one, it's going to be a little bit different. I'm going to show you right now. For example, you can make the evaluator model anything based on any criteria.
This is the magic of a Kapa loop. How can you get more agents running for longer periods of time without your involvement doing stuff on your behalf?
Get out email to score at least a nine out of 10 on the humanify cartridge. So, if you're sending out emails or newsletters to your list and you want them to sound human, you could put in something like a humanify cartridge where it analyzes the voice, the variety, the predictability, just like any AI slop detector, create an algorithm and give it a score out of 10.
And only once it ranks nine or 10 out of 10 is that email available to be sent by you. Another example, you might want your email to be extremely good on the marketing front. And of course, Alex Hormozi is a wizard with his marketing.
So, you could pull in Alex Hormozi's public tweets, his public emails, his public YouTube data, create a Hormozi AI algorithm that is going to score it out of 10 on how good that email is marketing-wise, and you can have it loop and keep creating a better and better email according to that cartridge until you reach a nine or a 10 out of 10. The last one we have, which is an interesting loop cuz it actually takes back real-world data, is an open rate email analytics engine where it takes your past emails, your past audience behavior, and your subject line patterns to try and predict based on real-world data that you've actually sent with your email list the predicted open rate and therefore also give it a score out of 10, which is the objective measure that the loop can grade itself on and keep looping until it hits that. In this example, we've actually put all three of these cartridges and all three of these evaluator goals together to get our email to score at least a 27 out of 30 from the cartridge evaluator. By the way, if you want to get these cartridges or build your own cartridges and build your own loops, come and join us inside the Dream Labs community. The link to that is in the description below. We are building cutting-edge AI systems for your business in there every single day.
So, you can come and use all of our templates or even create your own and learn how to do this and be on the cutting edge of AI tech. So, we've come back to Claude code to get this Kapathy loop out into the real world. We've typed {slash} goal because it's working towards a goal. Read the Claude prompt MD, then run the controlled newsletter improvement loop exactly as written.
Keep improving the local field notes signal draft until humanify Homos AI and open rate total at least 27 out of 30.
No cartridge can be below eight on the score. The body's under 150 words, the subject is under 45 characters, and there's only one call to action in the email. No block spam phrase appears.
Stop after five attempts. Save it as a draft mode only and never send it. So, we're going to send that off now to see how this loop goes. Okay, so that loop has finished, which running these loops is pretty incredible cuz you literally can walk away from your computer, you can start doing something else, and you know this email is getting sharper and sharper by all the criteria that I set it. And of course, you can plug anything in here from social media content, the websites, landing pages, speeds, products, customer service. Almost every part of our business can be looped. And you can see there's about five different attempts that this loop had on improving this email and getting it judged by those cartridges that we went through.
So, attempt three, attempt four, attempt five. Coming down here is a table where you can see the total score. It started at really weak at 7.1, and this was the original email. In today's fast-paced world, managing your inbox can be overwhelming. It's a very generic-looking thing. It looks 1.9 out of the humanify cartridge, 2.2 from the Homos AI, and the open rate was only a 3.0 on the rating, which gave it a 7.1 out of 30. And you can see even the buy now and the learn more, it's a very generic-feeling email. It then absolutely blasted. It almost tripled its score here to get a nine on the open rate, 7.5 on the Homos AI, and 8.7 on the humanify to pass, but it did not reach the 27 out of 30 that we wanted to. And so, it went again and improved it by another 1.0, which this is the updated email here which reads a lot more concise, a lot more personal than that original one. Then it hit number four, which is the best rating out of all of them, 26.8, 9.2 on the Humanify, 8.2 on the Humata AI, and the open rate at 9.2. Then it went on to five and actually fell back a little bit to 25.5. So, the best one did not quite reach our goal at 27. It was two points below it, but it was extremely close. And so, this took about 40 minutes to keep looping and get our email better and better all the way to attempt number four, which was this email right here. So, if you want to set up your own Kapathy loops and would like some personalized one-on-one help, come and join our Dream Labs community. We want to create the best AI community in the world, and I'm literally in there waiting for you to help create the perfect loops, the perfect agents, and the perfect omnipresent AI setup for your business. The link to that is in the description below. Thanks for watching. I'll see you in the next video.
Related Videos

TOP 15 Data compression Interview Questions and Answers 2019 Part-2 | Data compression | Wisdom jobs
wisdomjobs
281 views•2019-06-28

CTS 158: 802.11w Management Frame Protection
ClearToSend
4K views•2019-02-04

NDSS 2019 Send Hardest Problems My Way: Probabilistic Path Prioritization for Hybrid Fuzzing
NDSSSymposium
496 views•2019-04-02

How realistic is Cities: Skylines?
CityBeautiful
159K views•2019-02-14

GUIs & TUIs: Choosing a User Interface for Your Python Project | Real Python Podcast
realpython
2K views•2025-04-04

The OSI Model - Explained by Example
hnasr
225K views•2019-05-12

Cloud Computing - Introduction
elithecomputerguy
98K views•2019-10-07

From Traveler's Dilemma to Dynamic Routing | Demystifying Networking
IITBombayJuly
5K views•2019-08-04
Trending

Ben Crump dealt MAJOR BLOW after His Own Nolan Wells Autopsy FACT CHECKS him
DeVoryDarkins
50K views•2026-07-23

Gremlin Arrives… While Dorothy May Takes Another Step Forward
The-moons
10K views•2026-07-23

Trump War Chief SCREWS UP by Posting Video Leading Judge to ORDER an EXPLANATION!!!
LegalAFMTN
110K views•2026-07-23

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23