This project elegantly demonstrates how strategic reward shaping can refine brute-force exploration into superhuman precision. It serves as a brilliant bridge between abstract reinforcement learning theory and the nuanced demands of real-time competitive play.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
I Trained an AI to Beat Mario Kart Wii's Hardest Ghost
Added:You're watching an AI play Mario Kart [music] Wii. I've spent the last three months training it for one purpose, to beat the hardest ghost Nintendo ever [music] made, the expert staff ghost for Rainbow Road. Few have ever unlocked it, and even fewer have ever beaten it. But unfortunately, my AI just lost to it by 2/10 [music] of a second. Nobody taught it this track.
It learned the way you did by playing, failing, and trying again for over 1 million games. In theory, given enough time, it will only get faster. But in reality, this thing couldn't survive 5 seconds.
And getting from there to two attempts took every [music] one of those 3 months. It took obsession.
>> Oh, no. [sighs] No.
>> What?
>> No. No. No. No. No, no, no, no, no.
Don't do that.
>> It took many long nights reading research papers and many long nights debugging. To see how we got here, let's go back to the very [music] beginning.
So, what does the AI actually see? While not pixels, pixels are slow and honestly unnecessary.
Instead, every frame, the game gets compressed into 243 numbers. The first set of numbers describe the cart itself.
Things like its velocity, whether it's drifting, how charged the mini turbo is, and finally, you know, whether it's falling into space. And then comes the important part, which is the track itself. To see the track, we take 16 points along the center line ahead of the cart. [music] And at every point, we check four things. Where the road is, how wide it is, which way it bends, and whether there's actually ground there.
The last one matters the most because Rainbow Road's most famous feature is that it barely has any guardrails. There are so many places where you can just fall. In fact, there are more places you can fall than where you can drive. As for what the AI actually is, it's a neural network with under 3 million parameters. To put it into perspective, ChachiPT has hundreds of billions. So, by AI standards, this [music] thing is tiny, but it only has to do one thing, play Mario Kart. So, that's the loop.
The AI takes in those 243 numbers, runs them through the network, and presses button on a [music] controller every single frame. But to train it, I have to tell it what good [music] means and what bad means. It's tempting to do something like just reward speed or wheelies or [music] mini turbos, but all of those can be reward hacked. If I just reward speed, [music] I might not get a racer that finishes a race and just one that likes to drive into the wall really fast. Luckily, for very nerdy reasons, there is a theoretically perfect reward for what we want. Reward progress along the track and subtract a little every frame. With this [music] math, to get the best possible score means exactly one thing.
finishing the lap in the fewest frames possible. Let's see how this goes.
Oh, how >> So, it didn't go very well. We started from scratch with no help and no hints.
I hit train and I watch what happened.
It launched off the starting line, immediately turned left, and for no reason, just drove directly into space.
Again and again, I got to watch hundreds of copies of Funky Kong at the same time drive [music] into the same part of the wall for essentially weeks. It could be hard to tell if something about your setup is wrong or if you just need to be more patient. And in this case, I mistake patience for a wrong setup.
Because after the AI played literally one year of Mario Kart, it finished [music] zero races and it didn't even get close. I would go on to find out that I was punishing time too hard and so everything looked bad and the most optimal thing to do [music] was to end the run quickly. But I think there was actually a bigger problem. [music] The reward we mentioned is theoretically perfect, but it's way too patient. To finish Rainbow Road, you need something like 9,000 consecutive good decisions.
And this thing was out there asking questions like, "What happens if I fall off right here? Okay, what about here?
What about here?" Because, think about it, nothing in our reward tells it that [music] falling is bad. It would have to fall thousands of times and slowly work out that falling is not a shortcut. But we know that falling is bad and that's something we can safely assume. Sadly, that does mean we will never unlock an ultra shortcut with this. But I would rather do that than waste a million attempts figuring out that falling is bad. This is called reward shaping. We had small penalties that fence off the obviously bad ideas so that exploration gets spent where it actually matters. We ended up adding two penalties. The first one was for hitting walls. Walls kill your speed and so you get minus points when you do that. And the second penalty was falling. Falling kills everything.
And with those two penalties in place, our training actually stopped wandering and started climbing.
Let's take a look.
At first, the falls moved further and further along the lap and eventually we were [music] able to finish.
And then we were able to finish faster and faster.
I don't think we're ready for the expert ghost yet, but let's see how we do against the easier normal ghost.
>> [music] >> And just like that, we had our first win.
But I was still a long way from beating the expert staff ghost. To do that, I'd have to shave off about 20 seconds, which in a race is a long time.
Oh, [singing] I had a feeling that it would get there eventually, but I was impatient and I wanted to know, can we get there even faster?
[music] [singing] >> [singing] [singing] >> To find out, I had to analyze the actual world record run. Let's take a look.
There was one thing I kept hidden from the AI. The 2 minute and 25 seconds of the fastest Funky Kong and Flamer Runner run ever recorded on Rainbow Road. It would have been incredibly easy to cheat. I could show the AI every frame of the race and every button the player pressed and have it learn to copy the human. But honestly, where's the fun in that? I wanted this AI to learn from scratch with no human knowledge. So, I made a rule. The AI would never see this run. But nobody said that I couldn't look. I replayed the ghost and looked at every [music] interesting data point imaginable. Let's take a look at what it takes to have one of the fastest Mario Kart runs ever [music] recorded. First, let's look at wheelies. This run has about 58 of them, about one every 2 1/2 seconds.
This was way more than my current AI was doing. But we have to be careful not to reward wheelies too hard, or else our AI may not learn to wheelie when it actually makes sense. There's a delicate balance to all of this. Next, let's look at mini turbos. It looks like we have one every 7 seconds. Luckily, our AI was already good at that, but I think it was still interesting to see. I also tried punishing the AI for using more than one mushroom per lap, but I never quite got this to work out. Teaching AI to wait for a bigger reward later is one of the hardest open problems in reinforcement learning. It's like resisting quick dopamine to feel fulfilled later on, right? But there was one number that helped this run the most. After the opening seconds, this run's velocity never falls below 77. I decided that if our AI's velocity fell below that, we would punish it. There's something interesting I want you to notice here, though. If we wanted to find something like an ultra shortcut, which is clearly faster than any of these nonshortcut races, this reward would push away from it. As we add more reward shaping, we get further from what the true best run might be. But luckily, our goal here is to beat the expert [music] staff ghost.
Let's take a look at what we learned from this run. At this point, our rules to the AI are make progress, keep the speed high, create mini turbos constantly, use wheelies whenever the track allows it, don't fall, don't hit the wall. And to my surprise, this worked pretty well. I was quickly inching towards the expert staff ghost time and eventually lost by just two ten of a second. And then we had a breakthrough. To keep you entertained, I hired an esports commentator to narrate what you're about to see. So Josh's Funky Kong AI is locked in unloaded for takeoff [music] in this psychedelic arena. Oh, but did he miss the start boost there? That could cost him some vital speed. But now he can slide down the slope and build up vital momentum.
Oh, that's nice. Very nicely done. Some super somersaults. Funky is clearly having fun out there now to negotiate the bend in this first lap of three. If not, he'll definitely be missing out on some potential speed.
Oh, but he's nailed one there as he flies through the star tunnel. Some special moves on show now. All that training is paying off. Now he's motoring and it's magical to see. AND IT'S LIGHTS OUT. AND GO, GO, go as he glides into lap two. What's he got left in the tank now? He gets a boost down the slope and up he goes again.
Legs a Kimbo gets a smooth landing too. Now some tricky terrain to navigate through.
Over the holes he goes, not for the faint-hearted.
Twisting and turning towards this star tunnel again with its rainbow of colors.
But is the future bright for Funky Kong?
Can he beat the hardest Ghost ever put into a Mario Kart?
Nice control and another boost. He's really found his rhythm now and looks as cool as a cucumber.
More somersaults and he's into the final lap. Time to give everything he's got.
Look at the ghost. And now they're racing. It's on like Funky Kong. Did he get ahead there with the somersault?
This is intense racing.
And up he goes again. And there's the Mario Kart. It's ghosted in out of nowhere. This time he rides around the rim to keep his momentum. BUT THE GHOST IS AHEAD AGAIN. This one's going to the wire. Where's he going to land? Out of THE STAR TUNNEL. OH, THEY'RE NECK and neck now. This is racing of the highest order as they sprint to THE FINISH LINE.
OH, THERE'S THE GHOST BACK AHEAD BY A WHISKER with the line in sight and Funky Kong nudges ahead. IS THAT THE move that wins it? BUT WOW, [screaming] FUNKY PULLS OFF A WONDER WHEELIE of epic proportions TO SOMEHOW PULL IT OUT of the fire for the mother of all victories. INCREDIBLE scenes as Kong conquers to leave Josh jubilant.
[music] We did it.
And here's the thing, it didn't stop.
While I was editing this video, the AI kept training. There were no new ideas, no new rewards, just more time.
As of right now, our time is 234.163, [music] which is 10 and 12 seconds faster than the ghost we just beat.
Not too bad for something that spent its first year driving into the [music] same wall.
I want to give a huge thanks for the people that maintain the Koko repo because this project would not be possible without it. Or it would, but it'd be a lot slower.
Anyway, thank you for watching. Oh, and if you want to watch the full run start to finish, no cuts, stick around.
Oh, and by the way, I'm going to be doing AI safety research at Antropic.
Thank you. Bye. Enjoy. Subscribe. Love you.
Oh my god.
Heat. Heat.
A [music] heat.
>> [music] >> Yow!
Woohoo!
[screaming] Oh yeah.
Woohoo.
Yeah.
Woohoo.
Woohoo!
Woohoo!
WOOHOO!
[music] >> [music]
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

Gremlin Arrives… While Dorothy May Takes Another Step Forward
The-moons
10K views•2026-07-23

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23

FURIOUS Raskin CORNERS DOJ over Trump DARK PAST!!!!
MeidasTouch
237K views•2026-07-23