Stephen Blum effectively simplifies complex reinforcement learning concepts by applying them to a relatable project, making advanced AI accessible to practical developers. It is a well-structured guide that balances technical depth with clear, hands-on implementation.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Teaching AI to Play Game: Part 1
Added:Hey everyone, welcome on in. We're going to be starting a new adventure today.
Building an AI that can play video games. Well, at le we we found a Dino Runner game. You know, the one that when you the the like the web, what is it called? The Easter egg. The Easter egg in Chrome web browser that has the dinosaur running game. It's the dino run game. Hey, quick fix, you made it first. All right, made it on in.
Welcome on in. Hope you had a good uh Tuesday so far. We're going to be getting on with our new our new adventure here. We're going to have I know there's more we could have done on the YouTube app, of course, right?
There's like a lot more that we could have done and it could have just kept going and going and going. I was happy that we got the MVP, we did all the coding, we got the video, we got the streaming. Hey, Nea. Yo, I was thinking about you because I need your I need your advice. I need your advice. We're going to be training an AI to play a game. And this will be sort of reinforcement learning. So, I've never really done this at this level before.
I'm going I've made some decisions.
Reinforcement learning. Joseph. Joseph.
Yep. Exactly. Some reinforcement learning. I want to do this based on the data of the game.
Like, obviously, we could take pictures of the game, right? We could take pictures of the game. Let me let me let me pull the game up here. Open index.html. HTML. Here we go. So, this is the game right here. Paste. Here we go. So, you know, you know this game, the Dino Run game, this one, right? So, uh I want to do for for example, just some um reinforcement learning.
>> Serious uh going serious mode. Exactly.
Yeah, that's what we're going to do.
Serious mode. Almost done with my hardware lock. Hardware crack. Python wrapper. What are you doing? What are you building? Oh, lost that one. So, yeah. So, I'm going to We got to do a few things in order to make this work.
We got to in order to make this work, we got to do a few things.
>> All right. Let me pull the terminal window over here and then bring this I'll pull this uh Do we want to bring that down? Yeah, I suppose we could do that. All right. So, that way we've got everything in the right spot. And then I will bring this window down so we can see the whole game. And I'll probably make this a little bit smaller. Is that smallest we can go? Okay, that's the smallest we can go. All right, here we go.
Uh, and then we will stick. There we go.
Perfect. All right. Hey, Alex. How's it going? Welcome on in. So, uh, quick fix.
What What are you What are you doing?
What's going on? Uh, images are more fun. Alpha Go uses images. We could We could do images. I mean we could do a CNN and also it would be a lot easier.
It it would be way easier to deal with that using uh so you can solve using a very very small network if you don't use images. Yes. So I was thinking that we could either do we could take like screenshots of the image, right? We could take screenshots of the image.
Thanks for the hearts you guys. Or we could uh grab the data of the game and we could uh just train on that. So like there's like a lot of options available to us there.
So I'm not sure what the best approach is. Stream cord is in pirate English.
Wait, what? What do you mean? What do you mean? Stream cord is in pirate English. Oh, hey. Whoa, whoa, whoa. Uh, let me pull that over there. There we go. Okay.
Uh, live teaching AI to play game. Looks good to me. It looks good to me.
How to change your bots's language. R.
Okay.
All right. Sure. If it was actually solving the issue would take the actual data and make smallest network. Okay.
Yeah. So, that's what I was thinking.
Let me close this. Bring this back over here. Okay. So, yeah, we've got our we've got our game right here. I'd like it if on So, obviously when I lo when I leave the um what do you call it? When I lose focus on the window, it pauses the game. So I need to fix that really quick. So let's see where we get that by index. So today we've got we're basically doing setup.
We're doing setup and planning use language list. Oh, cool. I didn't know about that. CNN teaches more for learning in in in wood. Oh, wait. Okay.
Decide what the goal is and then which one becomes obvious. Ooh.
So we could so we could take screenshots, scale them down, right? Cuz I mean there's only so many things, right? And there's only there's there's up and down, right? There's up up and down. So there's like different moves that we could do. Hey Isaac, how's it going? Isaac edits 00. Welcome on in.
Happy happy Tuesday. Yes, I do. Yes, I do. So the goal is to play the game, the AI to play the game really well, right?
So, we want the AI to play the game really well. So, that's that's the plan.
That's what I want to do. AI plays the game really well. All right. So, focus.
Is there any sort of focus in here?
There's also what is it called when you leave the focus? Uh Google browser event leave uh window focus. What is that called? Uh I know there's I know there's a word for it. There's a there's a good word for it. Visibility change.
There we go. That's what it is.
Uh, what are you making right now? Oh, Isaac edits. We are making an AI that can train to play a game. We're training an AI to play a game basically. And this is going to be it's going to be the Dino Run game. How's it going there, Benjamin? Hey, Benjamin Crocker. Hey, set up. How's it going? Good to see you.
You are on uh Potato. What is this?
Potato dear stream. What is that? Oh, okay. Someone else's stream. Got it.
Hey, Zack. Thank you for the flow. Stay.
All right. My gratitude to you, Zach.
How's it going? Hope you're doing good over there. It's pretty cool, Isaac. I'm really excited about And you wanted a sleep machine, but he kept ignoring me.
Oh, no. That's terrible. That's awful when it when you have to be ignored like that. You don't use AI, use LLM. They're better for c for the country technically. Yeah, you're right. You're right. They're eclair cream. Good to have you here.
All right, we'll be building reinforcement learning. Yes, exactly.
That's what we're doing. We're doing reinforcement learning right now. We're We're going to be building it right now.
I'm playing the game, obviously. Right.
Can you react to Sleepy Machine? Can you react to Sleepy's machine? What is that?
I don't even know what that is. What are you talking about? Hey, Peter Parker.
Good to see you. Are we bringing dinosaurs back? Yeah, we are. Yep, we are. You may want to use JS today so we can run the AI in the browser and send replace use the controls with AI. Yes.
Yes, that's the plan. I was thinking about um transmitting the data, Neva. I was thinking about taking all the data and events in the game and transmitting them. So like the the current state, the current closest element on the screen like all the obstacles you see here, the different cactuses and things like that.
I want to transmit that. So then that I can subscribe to it in Python and then make decisions about what moves to make.
Uh did you have braces? Yes, I did.
Eclair, I had extensive orthodontia. I did extensive orthodontia growing up.
Send all the data. Yes, exactly. Hey, there we go. You can tell. Yep. Yeah.
Extensive orthodontia for like uh five or 6 years. So it was a long time. It was a very long time. Uh, your plate is so expanded. Oh, yeah. Yeah. Yeah.
Thank you.
Yeah. Whoa. Whoa. Whoa. All right. So, this is the plan for today, you guys.
We're going to be playing this game. So, I want to get the visibility change going. And then I also want to do a few other things on this. So, let's see here. Let's hide. How am I going to deal with this right here? Uh, pull this over here.
All right. There we go. Perfect. Okay.
And then we'll bring this up. Ah, we could bring that down a smidge. Right there. Okay. So, here's our code. All right. Let's see if we can find it. So, visibility change. Oh, wait. Oh, so that didn't uh that didn't find it. Oh, interesting. Okay. Oh, this is not even the file. Is this it? Oh, index.js.
There we go.
Wait. JS. JS. JS. JS. I'm like, hold on.
Wait. That's not JavaScript.
Visibility change. There we go. All right, that can work. It's kind of funny. Since the browser limits visibility of ports, you may need a proxy. Oh, you redacted your message. Hey, good news. Uh, we can figure it out. We can figure it out.
Hey, Neon Java. Hey. Hey. Hey. Good to see you. Welcome on in Rudra. Hey, Rudra. Good to see you. It looks complicated. Yeah, it's going to be.
Hey, Tori. Tabricio.
Uh, Tor Torio.
Torio. Hey, how's it going? Never mind.
Dumb. I'm talking myself out of it. All right. No worries. All good. I'm doing good, Julian. Thank you for asking. Hope you're doing well over there as well.
All right, you guys. So, we're going to take this change of visibility here, and we're going to let the game continue to run. On change of v visibility, bind handling tabbing off the page. Uh, and then what do we want to do? On change visibility, thisbind.
Okay. Um, let's see here. Oh. Oh. Oh, I see. Okay. So, it allow us to resume visibility state. this stop otherwise reset and play.
Okay, so on change visibility. So I think we could just comment these things out here, right? So we we need the game to continuously play. So this is what we need to do to see if we can make it continuous play. Hey Derek, welcome on in. Good to have you here. Happy Tuesday. Uh how you doing? You got the MacBook bit 48 GB price was like out of your range. Oh, hey Lily love. Hey, how's it going? Good to see you. Welcome on in. Happy Tuesday. All right. So, let's see if we could So, yeah, if I if I All right, perfect. Now, when we crash, I want the game to reload. I want the game to reload as well. So, we took we took care of this. So, where's is it crash? Yeah.
Okay. So, it's called crashed. So, when the game crashes, I need to set like a timer so it auto reloads.
Want a cookie? Oh, yeah. I like cookies.
What's the design for this, Derek? Uh, we're going to be, so far the plan is, let me see. You know what? We should probably write the plan. Let's write the plan here. Is there a read me file?
There is. Okay. So, this is the T-Rex runner game, right? The T-Rex runner game. And we're going to be converting it to have allow an AI to play it. So, let's see here. Uh, T-Rex Runner game.
AI learns to play. Here we go. So, we're going to do this right there. Hey, how's it going there, Plump? Hey, Plump Orp AI viewer streambo.com. Okay.
Advertisement would be ideal if you want to run the game headless and oh, 20x of speed. Yes.
All right. So, that's also something I was curious about. We could probably do that as well. We could probably do that as well. Stopped redacting my messages.
Wait, what' you say? What were you saying over there? Let me see. What' you say? Uh, search it on YouTube. What' you say? Okay. Um, if your message doesn't show up, it's because it gets automoderated. It gets automoderated.
I'll subscribe to you, Steven Blum. Yo, Lily. Oh, Lil Snow. Thank you, Lily loves snow. Thank you. I appreciate you.
All right. All right. Uh, Joseph, Joseph, what is sleepy? Yeah, exactly.
Exactly. What you could do is um Am Amanda Amanda send it in Discord. Send the link in Discord and we can take a look at it there. Send the link in Discord to your home address general chat. Yeah, exactly. All right. So, we're going to do AI learn to play. So, Derek, to answer your question, you don't use Discord? Oh, no. Really? Oh, no. Okay.
So, one second. That's the best way to send a link. Hey, Stephen. Not the BG score in your video. This is called uh actually this where is it here? Um let me close that. Here we go. Welcome again by Jenlu. Welcoming bin by Jenlu.
I'll put it over here in our Discord channel. Here we go. I'll I'll put this in Discord so you guys have it.
Go to Discord under link share. Hey, some cookies. I saw some cookies there.
Look at that. We got some cookies.
Check general. I just did. There we go.
Let's see. So, just search it up on YouTube, please. What? What is it? What do I What do I even mean? Search, please. Pin. Let's see. What' you say?
Check general. Pin the general. Wait.
Pin what?
What are you thinking of pinning over there? All right. One second. Let's get the game going. Uh, I Let's see. The plan is so some sort of DQN uh sort of re reinforcement reinforcement learning. One second. Let's spell this correctly.
All right. Did we spell that correctly?
Uh, okay. I I mixed up the reinforcement part. Here we go. All right.
There we go. All right. So, reinforcement learning. This is the plan. We're going to have the AI teach.
Your link got censored. Oh, yeah. Yeah.
Yeah, that makes sense. If I don't use Discord, you can at least search up Sleepy's machine on YouTube. Okay, we'll do that real quick. All right, are you ready, Amanda? Are you ready? Everything with this gets deleted. Yeah, it would.
Yo, Robot.
Yo, how's it going, Robot Blot? Good to see you. Welcome on back. Happy Tuesday.
All right, so what was it? You said sleepy time something here. Uh, I'm ready, Amanda. All right. So, you said sleep.
Where is it at? And if it's if it's a video, uh, I might not be able to play it. Oh, it's on YouTube. Okay. I'll we'll look at it, but I won't play it cuz of, uh, copyright.
We won't play it, but I I'll take a look at it. Let's see here. All right. Oh, wa. Okay. Sleepy sound. White noise machine. Okay. Noise machine surrounds you with a smooth soothing blanket of sound. All right. Sleepy Machine.
Is it Is it safe to play on stream? Do you guys know? Is it copyright free?
Don't uh puts us to sleep. Okay. All right. I won't. I won't. No, it's not copyright free. Okay. I can't play it.
That's a problem. We won't be able to play it. But at least we know what it is now. We know what it is now. All right.
Good to see. Good to see.
I don't think that's what she meant, Stephen. Oh, it's not. Uh, donut punks sleepy. Oh, was that a different one?
Okay. So, is that is there something else? Some other other thing going on there.
Okay.
Uh, I didn't get it. I didn't understand the I didn't understand. Tell me what I missed. It's like a YouTuber name. Oh, got it. Okay.
Punk. Okay, got it. Sleepy from Donut Punks. Okay. Sleepy from Donut Punks YouTube. Okay, let's take a look here.
Oh, here we go. All right, there we go.
Donut soda pop punk. Is that it? How's it going there, Futa? Welcome on in.
Good to have you here. Also, Sesame meet. Hey, Sesame meet. Which algorithm are you aiming for? I don't know yet.
This is what I need your guys' help on.
This is something. This is reinforcement learning. I've never done this before.
So, we're going to figure out how exactly to do it. We're going to see if we can train the AI to be successful.
What' you say? Uh, Mero, you confused?
Yes, definitely. Uh, you got to explain the details. You got to explain the details. All right. So, we're going to do reinforcement learning. Uh, and also when the game when the game fails, I want it to automatically start over. So, we need to figure out how to do that so that way it can, you know, give us continuous data. Google it without YouTube. All right, we'll try it one more time. One more time. Okay, you said sleepy from Donut Punk. P U N KS. Okay, there you go. All right.
Oh. All right. Uh, a What is this? I don't know what's going on. It's a fastmoving pajama clad character doughnut punks who attacks rivals through his grandmother's pills.
What? What's going on? Are you trying RL? Yes, I am. Um, yes, we are trying reinforcement learning. We're trying reinforcement learning images. Is this safe? Okay. Okay. Okay.
Yeah, it's pretty safe. It's safeish.
What? Yes, we're doing reinforcement learning first. You got to simulate the environment first. Yes, that's me. Yes.
So, what I need to do is I need to understand what the game is. I need to grab the output from the game. We could either we can do it via image, which could be a little bit more intense. Or we can use something called DQN. So, we're going to try that. There's Sleepy's machine. Okay. There we go. All right. We did it. We did it. Okay. Done.
We're going back. We're going back to coding now. What's happening? We're building a reinforcement learning as small as possible. Yes, that's what we're going to do. So that agent have all the actions and results. Yes. So there's an agent which is going to be the AI and then there's going to be the environment which is the game itself, right?
Uh you would anyway do DQN? That is the network name. All right. So, DQN is what we're going to learn about. What's happening? I'm confused. Uh, you got it, man. All right. Says me. Excellent. Good to hear.
It would be a CNN DQN. If that was the case, it would be a CNN DQN. All right.
Good to hear. So, we're going to we're going to take this Dino Runner game and we're going to train an AI to play it.
Right now, I'm playing it, right? I'm playing with my keyboard. I've got my I'm pressing the the spacebar button and that's what we're going to get. We're going to make this work.
So, let's see here. Reinforce losing DQN model, right? So, we're going to have a lot of to-do items. All right. So, make game auto restart uh without without user input.
Uh and then we need to um allow it. I I think that's the first thing we got to do here. Then we're going to need export data from the game.
So via uh p publish on channel. So that way we can capture the data and then we can if we see the data and we know how to like simulate the data then we can create a simulation environment outside of the game itself and then we can run it like 20 times fast, right? We can run it at 20x speed but CNN would take a lot of time. Yeah, CNN would take a lot of time. So, we're going to go just the the features just from the game itself. You work in an RO lab. You love to see people getting into machine learning. Hey. All right. Good to hear it. Good to hear it. Yeah. So, this is what we're going to try. We're going to try I've I've never done reinforcement learning at this scale before. So, these are the things that we're going to do here. We're going to do these two things here. So, that's the plan. So, let's make it so this can auto restart and then we're going to export the data. Those are the two things that we're going to do today. features would work better. Features. Yeah, there we go. Hey. All right, Luffy. Hey, Luffy fan. All right. Thank you for subscribing and joining the right channel for software engineering. Good to have you here today. I'm really excited about this cuz this is something that it always seemed really advanced and complex to me. After reading through some of it yesterday, it turns out we're going to be successful. I feel confident that we're going to be successful. Wow, the live caption is cool. I just saw you on Twitch. Hey. Hey. All right, Kobit.
Thank you. Thank you very much. Yes, this is technology I've written many years ago. Some tech that I wrote many years ago. Live captions, guys. Create a Python script that detects it and does it. Yes, that's what we're going to do.
Exactly. Quick fix. We're going to do Python for the AI side, JavaScript for the game itself.
Find a way to extract the game state from a C. Yes, exactly. So, I'm thinking like every 50 here. So, extract the game. So, uh this will be like every 50 milliseconds.
uh get game state publish to channel.
That's what we're going to do here. And then that will allow us to capture the game state every 50 milliseconds. So we'll just like snapshot it, right?
Can you can do it without AI? Uh well, I mean obviously I can play the game myself, right? Hey Redless Saber, good to see you. Welcome on in. Happy Tuesday. RL is the most pain to set up in the database and learning environment. So it's a lot of busy work.
I'm actually perfectly fine with that.
I'm perfectly fine with that. I will make it busy. So, sure. Playing the the no internet game. Yeah, that's what we're doing. Yeah, it's the no internet game. Remember sort of like that uh you know the Chrome browser when the internet's down.
Then you feed it into the algo. Yes, that's me. That's exactly. So, one of the things is I don't know exactly how the reward because I know there's like a reward and a punishment like when the game when you fail the game, it has a punishment and how you would apply that to the loss for the learning rate and the back propagation.
Voila, you made an AI agent. There we go. All right. Yeah, you'll crash.
You'll crush it. Nice. Thank you. It's not that complex. It's a bit hard and it requires some work. So, that's what we're going to do. We're going to get things going. We're going to get it going. All right. So, first thing I'd like is to make sure the game auto restarts. So, it's just constantly running.
And let's see if we can figure that out here. All right. So, this crash this crashed. I need to see when we assign this crash through here. Here we go.
Game over. All right. When there's a game over, we need it. We need to set a timer. Set time. Is it Is this Is this a thing that's here? There's not even a set timer. Okay. So, we'll say set timeout and then we'll pass in a function and then we will say after like uh I don't know 100 milliseconds.
Uh yeah, maybe maybe 200 milliseconds.
Yeah, sounds good. Hey, thank for the hearts you guys. Oh, hey Angel, Devil Potato, how's it going? Just joined.
What's going on? We're building reinforcement learning. We're going to be building an AI that learns how to play this game. We're going to learn it.
Didn't even know it was possible to play it without with the internet. Yeah, right. It's open source. This is an open source game.
The loss and bellman equation is very simple to compute. You can uh you can open some paper to explain it or I can write it up for you. Nea. Yeah, this is it. So, we're going to learn how to do it. We're going to learn how to do it.
Quick fix. Boom. Boom. You got it. Good.
All right. All right. So, let's see if we can get this to go here. So, how do we how do we restart the game? uh game panel sprite dimensions and then when when this when this crushed crash is called crashed for/c crashed equals so here's false restart. Okay, so we have to do a restart. So that's what we'll do. All right, so we're going to say restart. Wait, wait, wait. What was it called? Uh is it called this restart? I think it's this.reart is what we have to do. I'm pretty sure uh in order to get it to do Whoa, there's a lot of uh crashed equals falses and TRS here. Oh, box compare. Oh, so this collision detection. Here we go. Thus the collision detection out of it. You're going to sub. All right, Angel Devil Potato, thank you very much. Yo, Bonzupi, is that the Chrome game where it pops up when the internet is no workie? It is. Yeah, it is. You've seen some make an AI for this game. I think the channel was AI Warehouse if I'm not mistaken a long time ago. I'm looking forward to it. I'm looking forward to building it. All right, Stephen streaming over the internet while playing in the internet a no working game. Uh, some wizardry there. I see what you're saying. We got some wizardry. Yeah. All right. So, let me see. Let me figure out how to get this game going. So, I need to click I need to do restart on this. So, I'm pretty sure we just do like a this restart like that, right? I'm pretty sure that's it right there. I think that'll do the trick. Auto restart the game when crashed. All right, let's see if that works. All right, reload it. Okay, so it should restart on its own after 200 sec 200 milliseconds. Crap.
Almost. All right. Uh let's let's set that maybe to 500 milliseconds and then we'll put it down lower. Is there anything that returns? Get time step.
Okay, let's see here. Restart. Is it called restart? All right, reload. Okay, let's try again. All right, how about detect the obstacles?
We're going to create an observer. We're going to create an observer.
Wait, did you recreate it? No, no, no, no. This is some This is open source already. It's open source. All right.
So, that didn't do the trick. And also, restart might not have been We should probably look at the the uh the code here. All right. Oh, set timeout not defined. Okay. Well, that's the problem.
That's the problem. Set time. Here we go. Time timeout. There we go. All right. That'll do the trick. Much better. Much better. All right. Now, let's go over here. Close that. Reload.
Let's try again. Okay. Okay, let's see if it works. Let's see if it works. So, it's going to detect the obstacles by the game. Yeah, there we go. All right, now it's looping. It's looping all on its own. Perfect.
So, we're going to detect the OB obstacles by looking at the game state.
So, the game state is somewhere in this file. We have to find where the game state is, and then we're going to transmit it. That is a lot. There's three almost three 2,759 lines of code that we had to figure this out. Tangoon is open source code. Yeah, we are. Yeah, we are.
Strategy of jumping as early as possible for everything. You can compute that without anything fancy. Nice. So, that's what we're going to try to do. Look for pair ends.
I suppose we could do some obstacles.
So, see, we've already got we've already found the state right there. Hey, look at that. We got the obstacles on the screen. All we have to do is find the closest obstacle. Find the X position compared to it is cuz the dinosaur it stays it stays at the same spot on the screen every time. Right? So if the dinosaur is on the same spot of the screen, then we should be able to easily find the when when to jump and when not to jump, right? So that's what we have to do. It's an array. Nice. Oh, it's structurally different from Minecraft Observer.
I actually don't know. You're going to make some ramen. All right. Hey, ramen.
I like it. That will make you feel better. It make you feel better. Make you feel better. Right. Game state would contain a race. Yeah. So, we got the game state. So, we're going to grab the game state there. The whole code is open source. Yep. You can just download it, implement directly in the code if possible. Yeah, we could. Absolutely.
I'm going to extract I'm going to extract the state though directly so that way we can figure out what the state is and then we can create a simulation of the game because it looks pretty simple because I really I only see here, you know, jump state. There's a crouch state also, right? There's a crouch state and then the nearest obstacle. So that's all the information we really need. The nearest obstacle, whether it's a ground obstacle or an air obstacle cuz there's a pterodactyl, right? You remember the pterodactyls.
And then we need to know when specifically to jump. Also, there's a game speed cuz it gets faster, right?
Cuz you can see it's getting faster and faster as I make progress.
So we need to figure out how to add that in as well. So there's a lot of things that we got to do. All right, we're going to make progress here. I'm actually get maybe a bit more than the nearest obstacle. Oh, we could Oh, we Okay, so we could include more. Also, Nea, I was thinking, do we want to include previous states as the input, right? So, it's got like more of the picture or do we only want to give it the current state, right? The current state. So, the model can play can can plan the jump. Do we need like future?
Do we need future and past data? Right.
So, we got pterodactyls here now, right?
Oh, it's getting it's getting faster, you guys. It's getting faster. You only need one state. Okay, just one state.
All right, sounds good to me. Sounds good to me. Hey, we're actually doing pretty good. Are we going to get any coding done?
Uh, the the the speed will increase every 100 points uh so that it can be used as a way to determine it. Okay, got it. Okay, it t Oh, I lost. It takes 17 million years to beat the dinosaur game. probably because it's unlimited, right? It's an unlimited game.
All right. So, let's see here. Where how do we get how do we get jumping? Let's see here. Jump. There's like a jump state. Let's see if we can get this. All right. So, jump velocity, min jump height, jump keys. Okay. So, this is interesting. If jump. Okay. So, we got a few things to do here. Just a few task items on our list that we can take care of. All right, let's go back over here.
Click the play button. All right, perfect. Perfect. Hey, Rock and Brian.
Good to see you. Welcome on in. Happy Tuesday. Good to have you here.
It's not actually unlimited. It does have an end. Oh, wait. Chaos by Dom.
Really? This game has an end. So, now we have an objective. If we can train the AI to beat to be very good at the game, we'll be able to have it win and we'll be able to see the ending. That sounds like a really good goal. That sounds like a good goal to me. There's no hidden info here. Just previous state don't matter. All right. Good, good, good. I'm not an expert, but can you make an algorithm that calculates the distance between the obstacles and the dinosaurs? Have Yes. Yes, you can.
Redless saber. Yes, you're absolutely right. We could. I think it's more fun.
I mean, because it's such a simple a simple game. I wanted to keep this simple for reinforcement learning purposes. I'm seeing some JavaScript here. Yep, you got it. That's the actual game code. We're going to modify the game code. You do need to have the up down speed somewhere. Yes, we will.
We'll be grabbing that. Your patch of the game so it runs on the Zen browser.
Ooh. All right.
Fair enough. Sounds good. Hey, how's it going there? Uh, Cholan. Hey Cholan 95, can you make an algorithm that can make an algorithm? Yes. I mean tech that's technically what we're doing. We're going to be building an algorithm that teaches an AI, which is another algorithm itself. So that's technically what we're doing.
That's technically what we're doing. All right. So if it jumps, I want to be able to make it jump. I just want to see how this works. How do we make it jump?
Runner code jump. E-code. If this class container snackar show, what's that?
What's that? All right. So, on key down.
All right. E. So, our E is going to be our jump code. All right.
And not crashed then. Or touch start.
Oh, hey, here we go. Here we go. Okay.
So, this is what we can do right here.
This is how we're going to get the AI to make its moves. We can make it We can add it add it jump here. We can add jumping here. So, we can do the jumping here. You patched it. All right. Quick fix. Nice.
Where did you download it from the code?
Uh, let's see here. Uh, GitHub config. Here you go. Oh, wait.
Where'd it go? Oh, I cloned it. So, I did. I cloned it. I cloned it from somewhere. Here we go. Let me see here.
Is this it? Here. No, that was a YouTube. How about this? Is this it? No, this is a I was looking about reinforcement learning libraries. Is this it?
No, that's the DLSR stable baseline 3 that I was looking at the other day. Uh oh, is this it? I know it's around. It's it's it's open source. It's open source, so you should be able to Google it.
Let's see. Where is it at here? Uh not webgu. I was looking there's theqin. Um yeah, it's around here somewhere. It's around here somewhere.
Uh I mean, uh AI predictive programming.
Yes, that's exactly what we're doing.
Yes. Angel dead potato to go. All right.
Ooh. Oh, wait. You would have to code it for points to higher faster because you get the dinosaur to jump. Auto jump 7.
It's We're We'll We'll get the dinosaur to jump. We will. We will. We'll make it jump.
So, if you just Google it here, we'll just try this. All right. So, uh T-Rex Rex uh Rex running game open source here. This is it right here. I'll put it on link share. I'll put it on link share on our Discord. Link share. There you go. So, here's the source code to the actual game. That's what we're doing right now. There we go. Hey, 57 Rain.
Love your eyes. Well, thank you. All right. Proxy sensing. I don't know. Si simulations. Hey, Code Meitsu. How's it going?
Good to see you, Code Meu.
One day I want to see on stream an RL for coding machine code. Oh, what the guys are doing at Google. Hey.
All right. I like to see it. Oh, there is no way. Thanks, man. Absolutely. You got it. Hey, ported Jaguar. Hey, good to see you. Happy Tuesday. Good to have you here. All right, you guys. All right, let's see if we can get uh which language is this? It's JavaScript. This is JavaScript.
All right. So, if I miss your chat message, let me know. Hey, One Punch Man. Good to see you. No doxing today.
No doxing, One Punch Man. We don't want any more doxing. Is the dinosaur game sprite an actual or moving pixels? Uh, it's definitely a sprite. I I'm pretty sure, right? Cuz let's find out. Let's find out real quick. There's assets folder. Uh, open assets folder. Here we go. Let's take a look. Uh, uh, yeah, it's Here we go. There's the answer. It's a It's a sprite or Well, it's it's got like a what do you call it? It's a PNG that has a bunch of different run states, right? And it's the it's the wasteful way because the only thing that Well, yeah. Oh, well, actually, it's perfectly fine. It's fine. Yeah. And then there's some other states of the game here like some cactuses and things. Some pterodactyls, some clouds, things like that. So, I hope that answers your question. All right. Okay. One punch. All right.
Sounds good. Sounds good.
All right. So, let's see here. Oh, I would want to create a fork of this or something. Let me let me see here. Get status. All right. Get uh commit am added auto restart on game over.
I don't think I have a a repo for this yet. So, let's grab a repo. Yeah, we're going to make a new repo. So, we'll say new repository and then we're going to create one, Stephen. All right. So, wait, did we already do this? Did I already do this?
I don't think so. So, let's let's me double check GitHub Stephen O LB. Let me look at my repositories. I'm pretty sure I didn't.
I'm pretty sure I didn't. M uh large language evals. No, that's not it. No YouTube app. No. Okay. All right. T-Rex.
I I was pretty sure I created one T-Rex running game dino run re uh check linkure. All right, since it's packed in textures is probably WebGL. Oh, I didn't even know about I didn't even think I didn't even think so. You're not a coder, but could I build a bunch of apps in two days using AI? Yes, you can.
Mean, you sure can. It's actually learning. Love seeing that. Well, it's we we're at the beginning right now.
It's actually just it's just looping.
That's me. When it jumps, that's actually me pressing the jump button.
You can see it on on the screen right at the bottom of the screen. You see my keyboard here.
This is the devil's language, JavaScript.
Uh, can you add a patch of my index?
Well, uh, what do you do? What did you do to it? All All I need is just the game. So, here here's what I'll do.
Here's what I'll do. Uh, I'll create this I'll create this GitHub repository here. Dino Runner uh AI learning. Ah, that's a long that's a long, right?
T-Rex running game. Dino run AI uh RL learning. We'll do that. Okay. All right. reinforce segment learning AI model to play the dino run game. The t the T-Rex run. This is called the T-Rex running game. Okay. Public create.
Okay. And then we will copy and paste this here. And then we will boom. Okay.
Here, I'll I'll put it on Discord. So, if you want to contribute or see follow along, we can do that here. All right.
Perfect. All right. So, go over to Discord and link share. There we go.
There we go.
What are you doing? Game patch for the game without the Google Chrome. Check.
Oh, and simplified online only.
New Runner.
This is interesting. Okay. What does this do? What does that do specifically?
Hey, Sham. Good to see you. Welcome on in. Happy Tuesday.
All right.
Two days is a short time, but simple apps you can probably Yeah, you sure could. You can make an app in two days.
No problem.
No get ignore. Stephen, leave on the edge. Living on the edge. Yeah, you could see. You can see there's that was actually default. That was default. We can add a get ignore. So, we can say touch.get ignore. Ignore.
There we go.
Now it's added. Now it's added. Hey Obvious, good to see you. Welcome on in.
Happy Tuesday. You'll boost your stream.
Oh, Lunch Punch Man. Well, thank you.
Navigator lowercase Chrome. So, there's something in there saying Chrome.
Uh Chrome.
Chrome. What do you say? Navigator.
Uh vibrate on mobile devices. Nope.
uh platform.
Okay. Uh let's see. Mobile iOS is mobile.
What is exactly are you doing? That needs to be replaced with new runner.
New runner.
Okay. Document.load right here.
Uh else I don't understand this. This works on every browser. Like what what is the objective that you're trying to do? I'm head out. Bye. All right, Angel devil potato feel better. Enjoy. Enjoy the ramen. Enjoy the ramen.
Thank you for the hundreds. You appreciate it. Okay. So, I don't understand. Quick fix. You got to uh let me know what exactly you're trying to do there. Thank you. All right. Yeah, absolutely. Okay. So, let's keep going.
Uh let's see here. All right. So, now that we got the game to re restart, let's see. Uh auto restart. Let's see what else we want to do. All right. Make game auto restart. We're going to do a green checkbox here. We achieve that.
Good to go. Now we Let's see. Maybe I can make it just jump randomly. Let me do a little bit of random jumping just so that way it plays on its own a little bit.
You should be able to do a PR so that it's more obvious. Yes, that would be nice. That would be nice. I will not dox your stream. Well, thank you, One Punch Man. I appreciate it. Thank you very much. All right, let's see. Let's see if we can learn how to jump here. Jump.
Jump. Jump. Jump. Here we go. Okay. Show notification when the activation key is pressed. Is there a notification? I don't see one.
All right. So, if I delete this, it probably doesn't do anything, right?
Hold on. Let me double check. Uh, okay. Let's delete this.
See what that looks like.
All right, looks good to me. Yeah, see I don't even know what that code what that landing code did. What did it even do?
Uh, I will not dox your stream. Oh, well cheese co cheese kernel. Hey, well, I appreciate it. Thank you very much. Uh, very much. Thank you. Uh, listener key down function if key codes runner is jump. If it's a jump key code, this container element classless ad runner class snack bar show. Yo, Antonio Hamilton, Antonio Hamilton, thank you for subscribing. I appreciate it. Good to have you here. You join the right channel for software engineering. Today, we're doing some reinforcement learning.
We're getting we're just getting started. So, if you're interested to see how it kind of works, it's going to play out over the next week or so. Did you develop the code by yourself? This is an open source project. It's an open source project, so you can download this and run it locally. I'm modifying it so that way we can use it to train AI.
You'll not dock your stream. All right, good to hear it, Anime. Thank you.
Appreciate it.
Uh, sub sub broetics.
Let's see. Oh, Destruct Lotus. Hey, how's it going? When you're going back to Rust. Oh, I can learn. I can't stand JavaScript anymore. Oh. Oh, we'll we'll be going back to Rust. I want to do the Rust candle library. So uh that's a really so at destructive uh the next streams in next week maybe next weekish will be rust c maybe rust and the candle lib candle lib for more rust plus ml I think that's the plan join in the open source spirit there we go all right all right so let's see here so what is This even do this like some sort of magic something here. I don't know what the snack bar business is. Uh query it looks like icon disabled. H some sort of snack bar item there. What coding language uh do I use the most often? Uh Python followed by Rust followed by C followed by I don't know um I guess JavaScript probably some TypeScript in there. SQL that counts as a domain specific language. So, you know, we could just say that. Or we could just say that.
How about the DOM and Rust? Whoa. So, you're saying like HTML and Rust? You're saying that? HTML Rust. Wow. HTML Rust.
Wow. Wow. On stream. More JavaScript than C. Oh, yeah. That's true. That is true. More JS than C on stream. We've done C. We've done C. Okay. Let's see here. All right. Yo, I have a Yo, I have a Yo, thank you for the GG's.
I have a calico. Oh, a kitty cat. Yo, thank you for the GG's. My gratitudes.
Thank you so much. Good to have you here. I have a calico. Welcome on in.
Hi. Good to have you here. Thank you.
Thank you. Thank you very much for the GG's. I I like the good games. I like the good games. No problem. Hey. All right. All right. Thank you. Appreciate it. Yo, Nova for the 100. Thank you very much. appreciate you. Yo, Nova Scientific. Ah, I like my Nova right there. I love my Nova videos. I like my Nova videos. Thank you, Nova, for the 100. HTML stuff is in Rust usually performs worse than JavaScript because sending Wom slower than the the JavaScript. Interesting. It's not a lot of use cases. All right. Good to know.
Good to know. All right. So, what what do we want to do here? I want to keep looking for the jump here. So target key down this stopplay playing sound. When when do we get jumping happening? All right. So if we're not jumping and we're not ducking, play the sound effect and start the game for the first time. I don't need to worry about that. Okay, here's a jump code. Is jump key.
This is running and jump key. Then we say so is jump key is running this Rex end jump. Yo, peace lord. Good to see you. Hello. Hello. Hello. Welcome on in.
Happy Tuesday. All right. So, we're I want to make it I want to make We're going to see if we can make it jump.
We're going to see if we can like just make it jump constantly or jump randomly. Right.
How is my two alt accounts saying? I will not dox your stream. All right. All right. All right. Sounds good. Sounds good. Yo, I have a calico. Yo, thank you for let them cook. Appreciate it. Whoa, that's a big one.
W streamer, thank you. I have a calico.
My gratitudes to you. Wow. I have Thank you. Explanation point. Party poppers.
Copy. Face face. Thank you very much. I have a calico. We want to stream in the leptose or the dioxysis. What is What are those things? What's going on? What are you doing over there? All right. I want to make it jump. End jump. Can we do other kinds of jump? I want to make I want to make this I want to make it jump.
I want to make it jump. Jump key. This pause uh reset play. No. All right. So, let's see here.
All right. is jump key. So, there's one thing here that does that. Okay. Jump code. Jump code. Jump code. Here's another one.
Uh, that's not this one. Show notification when activation. I'm going to delete this cuz I don't know what that is. Just going to delete that. You know what? Instead of fully deleting it, I'm just going to comment it out here cuz I know this doesn't do anything here. Let's get rid of that. Okay, that's gone.
You're living in a bubble.
Hey, Snot. Good to see you. Also, I I love Calico. Thank you again. Appreciate you very much. Appreciate Appreciate you. Letos and Dioxus are Rust WOM FRONT AND OH, OKAY. Both perform worse than React on Startup. All right, good to know. You're actually underrated. Well, thank you very much. I love I love a calico or you have a calico. You have a calico. You have one. You have a cat.
Hey, Sergio. Good to see you. Welcome on in. Happy Tuesday. Dioxis is like really really bad. Okay, good to know. You have two alt accounts. Texting will not dox you. All right. Sounds good. One punch.
I really do have a calico. You do. You have a real calico. All right.
All right. I've I've uh I've had I grew up with cats. I grew up with cats. So, let's reload this. Make sure that works there. Okay. It should work successfully without any problems. Okay, it looks good. All right. Good, good, good.
Leptos is pretty good, though. Okay, good to know.
All right, now I need to figure out how I can get the jump going. So, here's one way that I feel like we could add in a jump command. So, right here, to-do, add AI jump command. I think we could do that here.
So, that way we can capture AI jumping the command here. So, if it's not playing, wait, hold on. Uh, let's see here. Target detail button. If it's not crashed and it's a jump code or uh it is a touch event, right? A touch event, like a screen touch event.
Uh, used to have a tuxed tuxedo cat and his girlfriend was a calico. Oh, okay.
Nice. Very nice. I have a calico. You got some you've you the cats. They're in your life. You've got multiple cats. Yo, Mark Lemon, good to see you. Welcome on in. Happy Tuesday. Good to have you here. Mark Lemon, good to see you. All right. So, here's another jump key. All right. Is jump key. Uh is jump key and it's running. Then this is weird. Like if the game is running, what is is running? Uh let's see here.
whether the game is running. Okay. Uh and jump key T-Rex. Jump. T-Rex. The jump, which is very odd because jump is complete. Falling down.
So, is this the way to start a jump?
How do we do that? I guess we could try it out. So, reached maxed height. End jump. Is there a start jump?
There is an update jump which is start jump. Here we go. So there's a start jump command here. All right. So if we're not dunking, if we're not ducking and we're not jumping, then we can start the jump.
And where is this at here? Oh, it was in this code all along. Okay, so we can start the jump. All right. So, I think what I want to do is add a an interval to like randomly try to see if we can get this to to get this to jump here. So, let's see if we can make it.
Yo, gloomy traffic. Hey, how's it going?
We're doing some reinforcement learning today. Good to have you here.
Welcome on in. Thank you for the moderation. Gloomy traffic. Appreciate you. All right.
In jump means it's done jumping. It's going down. Yes, that makes sense.
You're making code for the dino game.
Yes, I have a calico. What we're going to do is we're going to update the code for the dino game here. So, that way we can give it information to the AI to learn. So, we're going to have the AI learn. It's going to understand the game state and it's going to learn how to play the game really well. And the goal is I want it to play it so perfectly that we make it to the end of the game.
Want to make it to the end of the game.
Oh, how's it going there? Uh, Vulov. Uh, Vul Lava. Hey, Vocava. All right, G. This is a JavaScript. This is JavaScript here. JavaScript.
How are you doing? We're doing good, thank you. 21 cats. I have a calico.
That's a lot of cats.
Cool idea. I'm looking forward to it. I think it's pretty cool.
Maybe that's why I sent Let him cook.
Yeah, there we go. All right, so I'd like to be able to start the jump. Let's see here. Uh, let's see. This T-Rex is a what? Let me see what this is. When are we when do we set this?
T-Rex equals null new T-Rex. Okay, so we could probably set interval here.
Set interval a function.
There we go. Every we'll say every 500 milliseconds we'll start the jump.
Uh, this T-Rex start jump. Start jump. There we go.
Let's see if that does the trick. Let's see if that works or if it fails completely.
Uh, well, that uh Oh, it's just going to go forever now.
Oh, we probably we probably broke it. We probably broke it. That's why we probably broke it.
What's up, man? Hey, how's it going there, Crash Crackger? Good to see you.
Welcome on in. What's up? Hey, good to have you here. We're building a game.
Well, we we've got the game. We got the T-Rex game going right now. And we're going to train AI to play it. So, we're doing reinforcement learning today.
Do some googling. Leptos is faster than we act.
Well, you know what? We could we could try that out ourselves. We could destructive lotus. That actually is a really good idea. Let's see if we can make that work here. Uh I want to go to stream ideas. We can compare new stream idea. Compare leptos and js uh slash react.
See which one is faster, Womtos or JS React? And I would think it's hard to tell. It's hard to tell. This is JavaScript post. All right, got a new stream idea.
It crashed. You said crash. Oh, Cash Cash Crackger. Hey, Cash Crager. Sorry about that. My apologies. My apologies.
Cash Crackger.
You should try to make it go 5x. It's We can We can Okay, let's do that. Let's do that. Let's add it in. Let's try to make it go 5x. No, no, no. Here we go. Like that. Okay.
All right. Let's make it go 5x, you said. So, there is a speed for speed. I know there's current speed.
Is the current speed? Yeah. So, let's see if we can make the current speed.
Here it is. current speed plus acceleration. Okay, here we'll make it go really fast. There you go.
Reload. Here we go. There you go. It's going real fast now. Now it's going Oh, it's going real fast. See, it's go. It's going to go way too fast now. We're not going to be able to do this.
Wom uh Wom can be fast. Yeah, I think so. A thousand times faster. All right, here we go.
It's going to go really fast now.
Really, really fast. Let's see what the Let's see what the current Let's see here. Um, where where's the thing at that we did before? Here we go.
Okay.
Current C current says current speed.
Current speed of the game. Here we go.
Okay. All right. So, max speed. Oh, there's a max speed setting here. So, we need to make the max speed go like really fast. And then we need to make the acceleration really fast.
Here we go. All right. Let's try that.
All right. Is that going to go really fast? Yeah, now it's really fast.
Look at that. See, it's going real fast now.
Hey, Nina. How's it going? Hello, Stephen. Have a nice day, everyone.
Thank you very much. Good to have you here. Nina, what are you using? UV mapping. We're going to be grabbing the game state directly. So we're we could do we could do um what do you call it?
Uh CNN, right? So we could do image based classification, image based reinforcement learning. I'm going to actually grab the game state so that way I can run it really really fast.
WM sucks to work with. Not Oh, it is. I liked I liked I had a good time with it running on WOM with Rust and targeting the WOM runtime. I thought it was pretty. Resm does suck. It is not the best. You're right. It's not the best for sure.
Things going fast and nothing's loading.
Yeah, it's going so fast now. But I like how the score is going up faster, faster, faster.
How much score will you get? I have no idea. I have no idea. Let's see if we can increase it even further. Let's Let's Let's 10x it. Let's 10x it. All right. Here. Let's reload. All right.
Try that.
It's so fast now. Look at that. We're We're like light speed. We're going light speed.
Having a VM makes it a bit Yeah. Uh right. Yeah, that's true. That is true.
So, based on this information, it doesn't seem like there's an end to the game. It's It does not seem like there's an ending.
That's why we have Modern Rust. Oo, it's the Flash. Yeah, I know, right? It is.
It's the Flash. What's your goal, Alex?
We're going to be building a reinforcement learning AI model that will be able to play the game. It'll be able to play the game. Oh, whoa. Look how high I could jump.
I could jump really high. Whoa.
All right. So, yeah. Um, we need to Oh, wait. There's a gravity. Interesting.
Look at all the constants. Look at all the constants of the game. This is crazy.
It beats the Flash movie. Hey, there you go.
Uh, it's pretty good. It's pretty fun.
Maybe I'll make a dino game. This uh base exploit high score and syncs the cloud. All right. Sounds good. Sounds good. Quick fix.
Score is going to stop loading pretty soon. I don't know. We're going to find out. So, you're pulling the game state instead of painting the world. Yes.
What's the source of truth? The game engine. Yeah. The game engine is the real environment, right? So, we've got the environment and the agent. The agent will be the AI that's going to be learning through reinforcement learning.
And then we're going to put it into the environment and test it out directly.
So, we're going to see if it actually works pretty well. 3D dino game, a 3D version. There we go. Nice.
Sending compiled code over the wire takes longer than source code. So, fighting a lost cause. I love it. I love what you're saying over there. Okay, so we got some time here. All right, I want to Let's Let's bring everything back to regular speed here.
Red, redo, red, redo, save, reload.
Okay. All right. So, now we've got the game running. It's running regular. And I've got a bug here. Unintentional, an intentional bug. So, that way the game could just kind of load there for a second. So, what I need to do is I need to figure out how to get the T-Rex to jump. I'd like to figure figure out how to get it to jump. So, start jump. Let's see this. Oh, we need the current speed.
That's why. Okay.
Okay. It should automatically cuz it's the game speed. Okay. It's the game speed. All right. So, I suppose we could do that. The only thing is I do Yeah. Okay. Here we go. So now, now it should work. Let's try it. Let's try it.
Reload. There it goes. All right. It's jumping on its own now. Dino jump. Now, all right. It's jumping. It's jumping.
It's It's basically RNG at this point.
What we should do is we should probably bop this down a little bit. Oh. Oh. Oh, it's it's working. It's working. Jump here. Let's go down to 200. Let's make it jump a little bit more frequently.
All right, here we go. All right, now it's really jumping.
Could you use the same tech to send someone to a course or exam?
Pull. Uh, I I assume so. I assume so. I assume that's the case. Could you use the same tech to send someone to a course or exam? I think the answer is yes. Though I can't promise because I don't fully understand. The cats are fighting again. Oh no. Hey man, keep going. You're doing really good. Have a blessed one. Thank you too, Alex.
Appreciate you.
Sprite Drinkinker. My opinion is uh Oh, the chipmunk named Alvin. I coded him onto your PC. Okay. All right. Spray drinker. Very nice. Very nice.
Oh, it's over. Okay.
Hading anything else to argue? Okay. All right. Let me see here. Uh, let's see.
Yeping. I'm just I'm catching up on the chats, you guys. Uh, sprite.
Interesting. Interesting choice.
Interesting choice.
Can you tell when it goes up or down?
Because Yes, I can. If yes, you can send the entire state to a Python runtime.
Yes, that's exactly what we're going to do, Neva. That's exactly what we're going to do. You're getting a dog next month. Yo. All right. Good to hear it.
Uh, that sounds like a thing you could do. Hey, Recreum. Hi. Good to see you.
Welcome on in. Happy Tuesday. All right.
So, now we've got the dinosaur auto jumping. It would be neat to see I No, we've got it. All right. So, here we go.
Let's go. This get status. Get add get commit. We did the what do you call it? Uh, auto jumping. Okay, push that domain.
Oh, wait. Did the commit finish? Oh, it did not finish. All right, commit. All right. Uh, auto jumping.
Okay, perfect.
Push to main. Perfect. All right, there we go. Recreum scripts, good to have you here. All right, so let's see. Now I need to figure out how to transmit the game state.
Let's open up the read me really quick.
Okay.
Make auto auto jump. We We got a lot of tasks. All right. And see, make see we we need to make it jump at a certain point. You know what? I suppose what we could do is now that we know how to do the jumping, we know we know where and when we can do this. So, we can simply have like a runner that receives state for the game and we can just manipulate the state directly there. So, we need to make a state input section of the game uh for the AI to make like I don't know play moves, play decisions. There we go.
Uh AI. All right. So, we need to make that. We need to export the data. So, let's take a look at all the data that's available to us.
Some people call it emotional support dog.
All right.
All right, some emotional support dog there. How are you dealing with the temporal information? Is the CNN seeing a single frame or are you adding something on top capturing over time?
So, gloomy traffic, what we're going to do is we're going to grab the game state. So, current speed. Here we go.
So, see how we got some game state here.
We got the current speed. We've got the running time.
Uh, we've got the game is over. The game is playing. The game is paused.
Uh those are states of the game.
Let's see here. And then we've got like is jumping.
Uh and then crashed. I think crashed is another one. Right. So we've got a bunch of game state.
We're going to capture that game state and then we're going to transmit it to a Python process running on our machine.
That will be the AI to capture the training data. What I think what we could do though is in order to train it rapidly.
We're going to need to we're going to need to do some extra work. We need to simulate it. So what I need to do is figure out what data can capture. What are you using to capture the training data? Uh see here's all the code here.
Here's the actual game. So I'm going to grab the state from the game and I'm going to transmit it. So there's a lot of in there's like there's there's like 2,700 lines of code. So, we need to figure out where all the game state is.
Hey, local admin. Good to see you. Is the AI the machine? Hello, Skynet.
Hello, Skynet. You got it. What do we got here? Make collision box. Yes. So, we need to grab the current objects on the screen. We need to grab the current speed of the game. And we need to grab some other state. And then we can grab that. We can transmit it as information as input features.
And then we need to create a reward system. Actually, that's something we got to write down here. We need to create a reward system. All right. Um, and then we need to Python Python script py t h o n script for AI model. We need to find the model. We need to do quite a few things here. Lots of things here.
We need to figure out uh MLP model. We need reward system. Reward system. I need to make a simulation so we can train it really fast. simulation and then I need a transmission or transmission function transmit outputs like game gameplay movements right output gameplay movements we got to do that uh let's see uh and then simulate it there yo uncle Dream 7 thank you very much for the sparkles my gratitude to you thank you thank you very touch. Yeah, because of Exactly. It's not. Yeah, exactly. Yep. If you got a dog, you got to for mandatory touching grass.
You might be able to get away with sampling less stuff and just saving the state.
We're I'm going to try. We're going to try. I think we can I think we'll be able to do that, too. Casino game might be a figure out through Oh, hey. Yeah, I see what you're saying there. Can casino game also be figured out through like you? Yes. Uh technically, although casino games is RNG, right? It's RNG.
Yes, you can make an AI, but you need a lot of data.
Wait, you have Mac OS? Yeah. Hey, Lumer.
Yes, you know it. You know it. You name one. You name Darwin right there. Mac OS. You got it, Uncle Dream. I am good.
Thank you again for the donation.
Good to have you here.
See, so this is simple. May be possible to have very very few examples and train it on the loop. Yes, I think so. I think so. We're going to have a lot of fun figuring it out though. Data best base of examples. Simulation is a So, right.
So, what we could do is we can actually just grab the live data and then do it 1x speed, right? So, we're doing current 1x speed, right? We can do 1x speed. I think I'd rather do like a 100x speed though and train it. So, I want to capture the data and see what that looks like. So let's transmit the data.
So we need to transmit the outputs uh and do a lot of things. So we made it auto jump. Looking good there. Yo, Uncle Dream with the silver star. My gratitude to you. Thank you very much for the donation. Appreciate you very much.
Hackintosh or MacBook. MacBook. MacBook.
Hey Nittton. Hey Nittton. Good to see you. MacBook. MacBook. The most productive shorts live so far. All right, good to hear it. Free BSB best youame output. I could do some youame- a so you can see a little bit more there.
A little bit more. A smidgen more there.
Can we just write a Python script? Check if there is an obstacle front and let auto jump with the AI stuff. SVR. Yes, I think so. Uncle Dream, take care, buddy.
All right, job shift. Meet you next time. Sounds good. Enjoy the job shift.
Okay, let's see if we can make our challenge go to the next level. So, let's see if we can transmit the data.
Also, we got 20 minutes. We got 20 minutes because then I've got my meeting starting here pretty soon. You don't like Apple? Quick fix. Oh, no. Apple's fine. A little expensive, maybe.
Uh, the game is playing the game. Stop.
I need to go outside and take another loss. It's playing a game. The game is playing itself right now.
How is random number direct as a saver? Virtual saver. So, how can you? Well, we don't have a saver. What's a saver? What is a saver?
I don't know what the saver is. What?
Insanely insanely expensive. Yeah. Mhm.
What is saver? What is that? More like necessary.
It could be.
All right. Let me see. probably going to work fairly well if there's some amount of planning. I I think we're I think we're on a good track. I think we're on a good track here. All right, so we got the auto jump. We got to make it auto start. Let's do the transmit. I want to transmit data. This is the part I want to do next. So every every 500 millisecond, let's get the game state.
Let's get the game state. Make state input section for the AI. Yes. All right. So, we need to do that as well.
And I think we could figure that out. I think we're servers. Oh, okay. Got it.
Servers.
Got it. Okay. So, let's read that again.
Uh la. Let's see. Server. Server. Game.
Where was it? Can we just write a Python script to check if there's an obstacle in front? No, I I read that one. I read that one. Saver game stop. So, we I I'm just using my my local hardware. I'm just using my local hardware. We don't need a server.
Stephen, is my capture beaten yet? Oh, quick fix. We could if I had if I if I I we could train an AI to learn it. We could train it and we could train an AI to learn it. Call it ghost capture. It's a good capture. It's a good capture.
Okay. Export data. All right. So, that's the plan right now. So, let's see. We need to see what the game current speed is. So, we've got a bunch of stuff here.
We got our obstacles. Oh, perfect. Okay, so we've got everything we need. We've got game over.
We've got playing whether the game is currently playing or not. So, let's try this. Let's try this. All right. So, we've got all our data.
We've got this. What is this here? So, this is the runner. And do we ever only have one runner at a time?
Does the runner restart?
Okay. So, if the runner is a new runner every time. So, let let's do some console logging here. How do we do this?
How do we do this? Um, let's make a new function.
I think we're going to create a transmitter.
Do it. Do it. Do it. Do Do what? Oh, make an AI to beat your Oh, hey, Musty Potter. Good to see you. Welcome on in.
Thanks for saying hi. Yes. Okay, quick fix. Let's add that as a stream idea. We can do that. I think that one's actually pretty straightforward. All right, so we could do that. All right, so make AI to defeat defeat.
Defeat new post here. Uh quick fixes capa, right? It's a it's a motionoriented capture capa, right?
All right. So this will be AI maybe rust or Python I don't know post. Okay with RL.
Wait, what do you mean? Hey Surf Tube, good to see you. Welcome on in. Happy Tuesday. Hey man, I found found you again. Oh, Alex. All right. I wanted to ask, are you going to be live tomorrow?
Yes, live every day. This week we're going live 7:30 to 9 Pacific Daylight Time. So 7:30 to 9:00. Uh, and usually we do 8:30 to 10ish over there. We go.
That's what we usually do. Hey, Cardinal. Good to see you, Mr. Blum.
We're doing reinforcement learning. Yes, we are. We're doing reinforcement learning. You got it. We're getting the game ready to go. There's a lot of pieces and components that we have to be prepared for for this game. So, we got to do a lot of work. We got to do a lot of work. All right. So, I would like to be able to capture the runner. Let me see. Runner. Runner. Runner.
the current runner. Oh, here we go.
So, if that's the case, do how do we get the current runner? H singleton. It is a singleton. Okay.
So, if that's the case, then we can do So, we already got an instance. That means we only need to initialize here. All right. So, we need to transmit game state. transmit game state.
And we're going to do that right here.
Okay. So, I need to pull in a a file.
How do we want to do this? Uh I I've got I've got a file. Where where did I store this at? It's a file that lets me I'm just trying to figure out where I did the most recent version of this. Oh, okay. I've got it.
Aren't neural networks supervised learning technique? We do need to implement pre-pretend like RL. So a neural network is a neural network. It and just how you train it is what matters. What data do you present to it? What information are you providing to it? When you're calculating the output, the loss calculation in a reinforcement world requires rewards. We have to calculate rewards. Yo, how's it going there? PMAD Kumar, thank you very much for subscribing and joining the right channel for software engineering.
We are currently going to be building a reinforcement learning AI that can play this game that you see on the screen.
Stephen, can you just ask how do you manage balancing work with your social life? Uh, work balance life ratio for you an innovator like yourself. So, I work so uh gloomy traffic. I work right between 12 and 14 hours a day, maybe a little bit longer. and I stream for like maybe about an hour or two per day. So, it's like a pretty small amount of my day and then I spend a little bit of time answering like the comments and things. Yo, Nishant OP, good to see you.
Welcome on in. Neur neural networks are supervised, but they can be unsupervised when custom reward is used. Yes. So, we need to do a custom reward as well. We need to do a reward system here. We need to do reward system. Yeah, we do. right there.
All right. So, I'd like to transmit this. Uh, let's see. Um, I want to open up the index html here. Where is our script at? Here we go. Index script. All right. So, we need another script. We're going to call this pub knob. And I need to grab I need to grab this. Where did I have that?
Where's my most recent? I think I could probably just grab it from npm directly, right? We can do an npm install pubnum.
Uh, so let me see here. Uh, let's see.
Get node pubnum stream ss pubnum ss is what I want. Pubnub sS npm. Let's grab that here. All right. Where is it at?
Perfect. Okay. We'll just do a little bit of a little bit of copy pasty. And it it's really just one file. It's this file right here that I want. So, I'm just going to grab this.
Here we go. Okay.
Just this file here. PB paste.
PubNob.js.
All right. Good. Good. I think we're good now. Let's see if it reloads without crashing.
All right. All right. Looks good. Let's open up this. Uh, let's see here. Uh, got a problem.
What did it say? There's an error here.
M. No, seems good to me. Okay.
Reload.
Does that look good? That seems fine.
Okay. Yeah, we're good to go. We're good to go. Okay. Unsupervised machine learning usually means model is mostly responsible for its own reward function.
We don't have a direct correct answer.
Yes, cuz we don't know what the actual answer is. We don't know what the answer is. We know what happens when it breaks or when it fails.
So, let's bring this back up here. There we go. Reload. Okay.
So, now we've got our pub knob. Let me open that file really quick. Uh, this one. Here we go. Okay. So, we've got our see window. Do we have that? Okay. So, we already have our pubnub instance here. All right. And we're going to transmit this.
How's it going there, Nishant? What's this jewels do? Oh, it's donations.
Those are donations.
Check stream ideas. Okay, jewels of YouTube. Yes, those are just donations.
Let me see here. Um, let me see.
Uh, all right. I checked it.
I checked it. What do you want me to do?
Remote desktop. Ooh, nice.
Good idea.
You also want to build like people should. I start learning programming. I recently became 30. Oh, nice. Is it worth learning? Yes, I agree. I think so. Yes. Always worth it. Always worth it in my opinion. It'll you'll learn new things that you never learned before.
Absolutely.
So worth it.
Okay, so let me see here. I think that looks good. All right. Oh, wait. We got 6 minutes left, you guys. I let me see if I can transmit any state. Let me see if I can transmit any state. Um I I will be able to stay a little bit longer because I have to wait for them to start the meeting. Uh so there might it might not be immediate meeting start.
We got time though. We got time though.
All right. So this is where we need to transmit here. All right. This pubnub equals pubnub nub. There we go. empty default everything this.pubnobub.publish publish and then I think that we could just say what is this? It is a setup and then we've got a bunch of data here. So we've got channel and message. Okay, so this is the game channel. This game channel equals state game state game state channel. Uh how about we do this? T-Rex Dino GameS State with with hyphens. All right. T-Rex Dino Game State. That'll be the That'll be the channel. Uh yeah, channel game channel.
Publish channel.
Channel channel. Thank you.
This game channel.
Thank you. And then we're going to message. message is going to be a bunch of stateful information here. And then I'd like to create a little a little game state receiver.
Let's make a little game state receiver that runs in my terminal. A little game state receiver that runs in my terminal so we can see the output of the data.
Game terminal viewer. We need that right there. A little. And we'll probably do it like in bash or something, right?
with Bash.
20 languages I need to know. The first C, C C++, JavaScript, React. You suck at the game. Hey, the official Frost. Hey, how's it going? Yes, we have it. Auto.
It's in playing auto mode right now.
It's playing auto mode.
I can kind of control it. I can kind of control it here. It's actually hard mode right now. Mstack, Express, React, and Node big in 2026. Yes, sir. I think so. You had fun with Rust. Fun language. Yes, I love Rust is a really good language. Rust is the language of choice here, you guys. Rust is a choice.
It is the only game far better. GTA 6. I pre-ordered it. Did you guys pre-order it? Yo, it's actually pretty cool. Yeah, we're going to Oh, the the official Frost we're going to be build. It's going to be reinforcement learning.
That's what we're going to do. We're going to have the AI learn how to play the game right now. It's just kind of random. It's kind of random. It's random right now. Uh, let's see. Oh, you haven't yet? It's a pretty good game that I'm really much looking forward to.
Only play Minecraft. Minecraft is good, too. Where are you from? I'm from Seattle, Washington, United States.
Seattle, Wah, USA.
That is where I am from. Yo, hey, Marshall. Hey, how's it going? I watch your streams on YouTube. All right, it's good to hear it. I was wondering what your thoughts are on game development. I already do backend with Python. Game development is a lot of fun. It is very very satisfying and it ends up taking a lot of your life. It makes the hours just like disappear. The hours disappear like they fully disappear.
I'm I'm like slightly manipulating. It's like hard mode. Yeah. Yeah. No, I lost.
Okay. Quantified quantum. Yo, wait.
Cozy. Hey, Cozy Quest. Yo, Quantifi Quantum. Cozy Quest CC. What is the deal? What is the deal? Hey, local admin. Hey. All right, local admin. Yo, is that you just used it? You are a new Wow. Nice.
Let's see. Rust. That is awesome.
Heart. Let me do that. We do a heart.
Okay, we do a little heart. Heart here.
One second. Wow. I was not expecting that. Hey, local admin. Very nice.
All right. Uh the official It's 9 in the morning. Oh, yeah. It's 9 in the morning. Oh yeah. See, I actually I have to get I have to I'm going to jump up on a meeting here on my phone and as soon as it as soon as it gets started then I have to end stream. I have to instream.
You're going to get me. All right.
Sounds good. Sounds good. Uh let's see here. Legend coding up a legend. Respect to Steven. Yo, the official Frost. Thank you. Appreciate it. That's really nice for you to say. Thank you very much.
All right. So, let's see here. I'm going to say test, right? Test. Actually, we could do this test. One, two, three. Oh, what we should do is um this dot there.
There's game state, right? There's game state here. Let's do let's do speed.
Isn't there a current speed? I think there's a current speed in here. Yes, there is. All right. Let's do current speed. All right. So, this is current speed. Speed. There we go. Like that.
No, no, no. We keep this here. All right, perfect. Perfect. All right, so now we're going to do game state.
Transmit game state. Uh, and we need to do this inside an interval. So, let's do this. I've got a better idea. Uh, set interval.
All right, we don't need to execute this right away. We'll throw this here. Every We'll do We'll do every 500 milliseconds to start. To start. Would you make a repo for this project? Yes. Indeed. In fact, we have one. Get status. Get add get commit.
Uh, started adding game state transmission, right? Transmission.
Push to the main. It is on our Discord.
Let's see. It's around here somewhere.
Let me see if I can find it real quick.
Here we go. Here it is. Yeah. And I believe it's under link share. Link share. Yep, there it is. It's already there under link shirt if you are interested. And there you go. Good to go. You think coding is still worth it?
Nice. Me too. Supposed to make any mistakes and code in to end even as much as the situation coding skills. Still learning is important. It is both due to financial constraints and the terms of development and innovation. You get a lot of innovation. You get some innovation there. All right. Hello Stephen. Hey Tugra. Good to see you.
Happy Tuesday. All right you guys. I have to jump on a meeting. I see now.
Looks like I'm ready to go. All right, I got to get going. One second. Uh, they say ready to go. Okay. Oh, wait. We should be up in the next 10 minutes.
Nice. Oh, no. I don't have to go. I don't have to go. I have 10 minutes. I have 10 minutes, you guys. Hey. Hey. We got some time. Bye, Stephen. Keep smiling. Gloomy traffic. Hey, I'm not going yet. I'm not going yet. I've got 10 minutes left.
Okay, we got 10 minutes. Yeah, this. So, I'll I'll be Let's see if we can keep this going here. Let's see if we can make this happen. All right. All right.
Hey, I thought I had to go. We will here in a little bit, but not right away. Not right away. All right. There we go. All right. Do you stream every day? I do.
Laxman. Laxman. And I sure do every day.
Every day. What I'm going to do is I'm going to get this running on my phone so I know when it's ready to go. One second, you guys. Uh, let me see here.
Okay, one second. I'm going to bull pull this up here. Yes. Finish the project you have in 10 minutes. No pressure. I know, right? Only 10 minutes.
Pause that. Okay. All right. So, now my phone is right there. My phone is right there. Okay. Perfect.
Okay. All right. DQN or PPQ? Oh, I don't know yet. I don't know. Hey, we we need to figure this out. So, MLP, we need to figure this out. So, we need to decide DQN or PPQ Algo. Which one do we use? Do you know? Do you know which one we're going to use? I don't know. I don't know which one we're going to use. We're going to find out, though.
We need We're trying to We're trying to modify the game really quick, though.
All right. So, I'd also like to add in a way to collect data on this channel. So, I do have a subscribe here. Oh. All right.
>> Hey, Stephen.
>> Hey.
>> Hey. How's it going?
All right, I do got to go now. All right, >> everyone gets three chairs.
>> All right, can you >> All right. No, I do got to go now. Oh, oh, I thought I had 10 minutes. All right. Sorry, guys. Thank you so much for joining today. We'll be back tomorrow, 7:30 a.m. Pacific time, and we will be continuing the reinforcement learning. Okay, have a good day. Thank you guys. Appreciate it. Had so much fun today. All right, let me do a quick little commit here.
One second. Okay. Yeah. All right.
Okay. Let me let me get everything wrapped up here.
Get status. Get commit am uh transmit game speed.
Push the main. All right. Perfect.
Minecraft has a monster been slain by player 30. Okay. All right. Thank you guys so much for joining. Good to have you here. We'll back tomorrow. Bye everybody. Have a good rest of your day.
See you tomorrow. Thank you for joining.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

MIC DROP: Smithsonian Director Called Out For Woke Propaganda
TheAmalaEkpunobi
37K views•2026-07-23

2.4 BILLION Records Got Leaked...
DeepHumor
15K views•2026-07-22

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23