The video offers a sobering reality check by distinguishing between mere computational efficiency and the nuanced wisdom inherent in human judgment. It correctly identifies that while robots can process data, they still lack the cultural soul required to navigate the complexities of our world.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Robots Are Already Here
Added:We're going to look at seven different ways to create AI images that have double meanings. They work in two ways at the same time. I mean, look, think about it. Zebra on a crosswalk, comb piano, single red cabbage sliced neatly in half. What is going on?
>> Let's do this.
>> Perfect. I'll say five right after. And we'll just keep it rolling until we get all the way there.
>> Yep. Exactly. So, I'll say five and we'll just keep that pattern going one after another until we reach 100.
>> Exactly. All right. I'm on it. So, I said five and you'll say six. And we'll just keep this rhythm going all the way until we hit 100. Let's do it.
>> Exactly. I'll say six.
>> Stop saying exactly.
>> Exactly.
>> Stop saying exactly.
>> Okay. You have to see later in the video if they make it to 100. Let's talk about whether you're noticing a certain design aesthetic happening all over the internet and if that could be Claude Code kind of standardizing everything that we view. This is um Allen's conservative countdown to AGI, but it's just stuck at 97. And yeah, I get it.
Like maybe a couple more times this year, he'll ratchet it up to 99 and then we'll get to 100. I don't know, maybe middle of next year or something, but it's going to be like this last 3% will be the whole embodied thing. So, it's once drones are like knocking at your door and talking to you when it kind of hit 100 here. We'll talk about a new paper that actually can teach models to get smarter and smarter without fooling them, without actually having an error rate where they're wrong. Dude, I'm like literally freaking out right now. I Okay, I'm going to take a little break here. All right, listen. I I don't know if I just caught something here on camera and I pulled up Instagram with it's a very unhealthy habit and look at what video is playing. I mean, the odds of this exact video, I mean, maybe a video of him wouldn't it be the craziest thing? I might see a video of him once every week or something, but the exact video I was just talking about shows up on my Instagram. And you think that it's not listening. That is such a coincidence, right? Am I tripping? I mean, maybe because am I I'm on the same account. I was just watching a video with an animated version of him, but why is my phone I don't know, dude. That just that's unsettling.
I don't know. Maybe it's explained because it's the same account and then it like related content. ood. Excited to introduce you to the sponsor of this video, Chat LLM by Abacus AI. This is an entire sandbox little universe where I am not logged into anything. It does not have low-level access to the computer that I use every day. It already includes all of the different models I could even dream to jump between. the current state-of-the-art GPT55 and Claude Opus 48 along with Gemini's best model 3.1 Pro, a bunch of various other models, and even a smart router. I mean, look at this. You can even deploy OpenClot to a supercomput. So, now you get the power without the risk. And it's not just the coding models. You also can generate images and video and speech.
So, Nano Banana Pro Seed Dance 2.0. Oh my gosh, look. Oh, is that a Grimlock?
Oh, sick. Please tell me that's a T-Rex transformer. What? That's you. So, you got your choice between the best models.
You can see VO over here or image editing with GPT plus Topz upscaling. It is all included in one place. And it has browser use because remember this is like its own little computer. So, I'll click try now.
Open LinkedIn. Find 20 CEOs of SVT tech companies and send a connection request.
Submit to agent. I get this interface which is super comfortable. It looks like just like everything I have with like Gemini and Claude and chat GPT chats on the left here. My projects and look at what we have here. So we have an entire computer with its own web browser. And look, my mouse is now inside of this computer. Like my personal data isn't in this computer. So you can get started for only $7 a month and then $10 monthly after that. And of course, lots more to explore potentially in the future. But if you found this tool useful and you sign up if you wouldn't mind using the referral code in the description below, I get $5 for each friend I refer. But um I really think it's a great product. This paper is called the physics informed neural network framework for elastodnamic wave propagation in bi material systems. But what it really means is that if you take any material like a big block of steel and you you know smash it with one side of a hammer, can a neural network predict a stress wave as it moves through the material. As AI models start actually making decisions in the real world, we need better ways to trust them. New AI research can improve how computers interpret the world around them. There's a new hybrid AI model that cuts financial forecasting error. Police in Western Australia have deployed live facial recognition technology. The police commissioner says it's not about mass surveillance. So maybe it's just about safety, which which is valid. It's just it could also easily be mass surveillance. You know, look at a softer conversation about literary translation and how it exposes the limits of AI. I mean, look, we've all seen Star Trek when there's a universal translator.
Like, that's a future we're supposed to have and I feel like we should be there.
After 7 years of working on a major league baseball AI to just detect if a ball went through a strike zone, and guess what? For how deep these billion parameter neural networks are, it might just be that one layer is enough. All right. What you guys think about Boston Dynamics putting a robot at the World Cup, not in the World Cup yet, but do you believe that just I don't know, few years from now that people might be entertained watching robots like this actually play soccer? I mean, do you ever see a stadium full of humans watching robots play? I guess there'd be no reason for a stadium full of robots, but like robots playing basketball, um, dancing. I kind of feel like there is a place for that because it's a spectacle.
It draws eyeballs. It gets people talking and, you know, the robotics companies want to compete and show off who's, you know, best and like they're putting billions of dollars into these robots, so they want to, you know, get attention. So, this could be the first of, you know, a robot at a World Cup. I mean, there was the NBA and then there was the WNBA. And why couldn't there be an RBA, right? Like a robot NBA or like an RNBA.
Maybe World Cup for humans is every four years, but then you stagger it every two years with, you know, robots from every country gets to submit the best robots that they have. Oh my god, that would definitely do. Imagine China would definitely be the favorite. Okay. So, if you're kind of an artist who's trying to just generate visually striking stuff and you don't want everything to look super generic, like people who don't really prompt much, there's a setting that you can just sort of tell something came from Nano Banana or from Chad GBT.
But if you want to craft a prompt, here are some cool tips. I mean, you can actually say in your prompt things like poetic, surreal, metaphor, make an image with a double meaning. Another tip Halfjourney Labs had when um they wrote this article was that you can take two things that sort of have the same shape and put them together. You can kind of imagine how a keyboard has this sort of similar feel to a comb. And here I was thinking if you guys ever look at our perfect fit, like Reddit perfect fit, it's always got stuff that like perfectly fits. See how like the microphone cables line up perfectly with the cable that hangs out from behind his TV or this dime in his ring. I mean, look at this. This guy's blood pressure medication fell right into the hole on his paper. Like perfect. Oh my gosh. So anyway, you can think like that when you're trying to put some images together. It's just visually striking.
Another trick is if you see an object, just just keep it there, but get in on it and have it sit awkwardly, right?
Like this is a thing like uh put like a sand timer on a floaty thing in a pool.
This is something that you might think about like it's like you're looking at it and you're noticing the sun going through the glass and you're wondering if there's a metaphor here or what the meaning is or like you're thinking about time and this being out in the sun using cultural reference the shadow technique.
I've seen this in a few paintings and stuff but it's just kind of cool. I mean clearly this pawn is like imagining could it one day become a queen and it all just comes through in this one photo. I feel like taking shadows and manipulating them into things is totally like a way to make a really cool kind of concept for someone. I mean, look at that. A dangerous peace symbol made with barb wire. A match that you know is about to ignite all the matches around it. I don't know. It's just fun to get creative with that kind of stuff. All right, I'll just play a quick little clip of this. Like, I don't even know why, but I just for some reason really found this funny.
>> Got it.
>> Yeah.
I'll say one and then the next day I can say two and we'll just go until we hit 100. Ready? I'll start. One.
>> Yep. I'm ready. So, I'll go after you.
Let's keep it smooth and we'll just count straight through. Go for it.
>> Perfect. So, I'll say one, you say two, and we'll just keep that rhythm going.
Ready? Here we go. One.
>> Exactly.
>> Perfect. I'm ready. Go ahead and say one, and I'll jump.
>> Okay. So, next up, I wanted to talk a little bit about talking to your phone like a person. So, I've been using Gemini Live quite a bit. And here are some things where I still find it awkward. It gives me sometimes pretty short answers like I'll say tell me all about like sleep no more or something and it'll kind of give me maybe I don't know 45 seconds of answer even though I might ask for a lot and then it asks me a question and the question's not like don't derail me. I'm trying to like solve a problem here. I'm trying to think through something. And then I don't know if you guys have this other issue too and I would have it with um Chad GBT a lot when I was in that speaking mode is it thinks I'm talking when I'm not. I'm just waiting listening to it and I guess it thinks I'm trying to interrupt it which is supposed to be natural like the way he does in this video but sometimes I don't talk like I don't know if it hears like something in my environment and it's like oh are you saying something I better stop talking or whatever but I find it pretty annoying like I just want you to finish your sentence. I'm sitting here listening. So that has to get more accurate.
I think in some cases the questions have to go. I don't like feeling pressured like I have to jump in the second it's done. I wish it just stopped and just waited, you know, or or I wish I could prompt it and say, "Hey, like if you keep it lively, I'm I'm kind of bored.
I'm in a listening mood." Like just keep talking. Or I could I don't know, signal it in some other way, but I just don't know. I think it's weird the timing thing needs its own AI level to it. it it is convenient and especially if I'm driving, I do think it's kind of fascinating, but I also don't want to be so like I feel like I need to talk to it because it's like waiting on me to say another question. And then also, I was wondering, do you I mean, I guess this will probably be a problem that will get solved, but do you think AIS know when they're in this kind of talking mode if another AI is talking to them? because it feels like these three robots when they recognize that all three of them are there should kind of be aware or like somehow they need to find a role in the background or to be part of a group.
Like if you say, "Hey, it's me and we've got three AIs here. Why are they all trying to answer?" Like as soon as they start talking over each other, maybe whichever one wasn't like the fastest should just wait and or maybe ask about the others like, "Hey guys, like we need to count to 100. Do you think um you guys cool with me leading the charge? Or if not, who else wants to do it? Or you know, like one of them proposed a plan, but the others didn't listen in a way where they didn't think it was a one-on-one conversation with this guy.
Anyway, I know I'm taking a joke and making it too much, but it's it's going to get solved. And we're in an interesting state where we get to see all these things like deployed out in the real world and then naturally, you know, end up with these issues and then solve them. And it's just funny. Gemini, count to 100.
>> 1 3 4 5 6 7 8 9 10. Would you like me to continue or did you just want the start?
>> No, that was perfect. Thanks for counting to 100.
>> You're welcome. Is there anything else I can help you with today?
>> You know what? I was going to show you that because I thought the problem had been solved and this was an old video, but it literally didn't say the number two and it didn't count to 100.
Not there yet. Next up, let's talk about the design aesthetics of the internet.
Um, I have recently updated my website and I've been looking at some people who are using Claude to update their websites and yeah, there's a sort of similarity to them I've noticed. And this article is arguing that there's a growing number of websites and sales decks that are starting to share the same visual style and Anthropics Claude and their design tool, their canvas tool is a big reason why. So, some designers that were interviewed by The New Yorker noticed different companies presenting nearly identical layouts, colors, typography, page structure, even though the companies had no connection to each other. So now there's this kind of if you don't make any adjustments, you get this cream background, rusty orange accents, oversized serif fonts, spaced out headings, rounded dashboard elements, and a handful of other design elements which seem to be Claude Design's default. You know, it's just interesting because it's not bad, but it's a lot. And at first, I was reading this article thinking like, oh no, everything's going to be so like bland and boring. But then I also kind of thought, well, you know, a ton of people were building websites on Squarespace and on templates for like WordPress. And yes, some really pushed it and got creative, but a lot of people were just like, h, yeah, that looks like professional. I'll just go with it. So, I'm not necessarily sure the internet was that creative to start with. But it does show that distinctive design is possible when you put in creative effort, and a lot of us just aren't. And because it makes you know quote good enough designs so easy that originality becomes optional. So in the future not only should we try to take the option to be unique and to put effort into our designs but also interfaces being generated through canvas are about to become everything which is a weird thought. I mean imagine if everybody's iPhone displayed their desktop or their icons however you wanted. Maybe you want one big button. Or maybe you want something smooth that has this kind of wave thing to it. Maybe you just want a bunch of text to click through, some like simple blue links. I don't know.
Like, why couldn't your homepage just be whatever works best for you? Maybe it looks like Doom and you like, you know, navigate left and right to find something fully 3D. I don't know.
Arguably, it would just change the context depending on what you're doing.
I mean, if it's like right before bed, maybe it just brings up the alarm. Or maybe if you're watching a show and then you need to like turn off an alarm. It just somehow like splits the screen automatically. Like it codes all of that just for you just in that moment. It's contextaware. And we have something in the near future that's like Fable but maybe you know a few times better even and thousands of times faster. And it just, you know, snap your finger like by the time you can click on it and observe it like a fraction of a second boom like tens of thousands of tokens built a custom web page or interface or whatever it is just or diffusion image like right for you right then. How crazy is that going to be? Like that seems very plausible in the near future to me. All right, next up let's talk about the nerdiest of all papers that I've covered for a while. physics informed neural network framework for what else?
Elastodnamic wave propagation.
Elastodnamic. I guess it's just like how much something wiggles. But there really is some fascinating stuff here that seems like only an AI could ever figure out. Okay, so imagine how important it is to know exactly what would happen to a building under all sorts of earthquake conditions. Imagine knowing exactly what's going to happen in an airplane if a bird or a drone hits it and how the steel would be impacted, how things would ripple, where the stress would go.
And you can imagine that to pretty much every product that we buy, every car, every gadget, every everything. So this paper is trying to figure out if a neural network can actually replace repeated physics simulations for waves that are moving through in this case steel and alum aluminum but it could be all sorts of materials in the near future. So a fast impact can send a stress wave through two joint materials steel and you know aluminum but it's really hard to predict what happens when that wave reaches the material's boundary. Some of the wave is going to pass through, some of it is going to be reflected back. So this study, they these researchers, they made a physicsinformed neural network. So it's called a pin model. And I have we talked about pin models before? I've I read about them once, but I don't think I've ever like really covered it. But they're physics informed, meaning the neural network was trained on all of these physics simulations, but then it's also trained sometimes through reinforcement learning on actual results and then it makes a prediction. So it's learned physics. It hasn't learned words in the same way like Chad GPT has or next token prediction. It's trained with physical rules literally built into it, not just data about physical rules, you know, like a model like uh AlphaFold that predicts the shape of something like that was I think it was from either from 1 to two or from uh AlphaFold 2 to three. That was the big change is like the the core of it was actually built around geometry, not walls of text about how things happen. And the researchers are able to compare this pin model, this AI model against actual uh what they call analysis fine element simulation.
So this would be something like a big Nvidia server that's doing an actual physical simulation. The model is predicting all sorts of physical things, axial motion, radial motion, wave arrival, peak response, and face averaged behavior. and it's having very strong agreement with the finite element simulators. Okay, now this is a big deal because remember an AI isn't perfect. It can hallucinate. It can be imperfect at sometimes, but it also can be so much faster than a physics simulation, especially because physics simulations, you know, in essence, you have to get almost as accurate as every atom in the material. or if you make some assumptions with algorithms, could those assumptions have weird little things that don't actually make it a true simulation? And in some cases, this pin model did really accurate comparisons.
In other places, it had weaknesses. They would appear after a main wave peak, especially in steel, didn't bounce off the joint correctly, where later reflections and stress changes were a little bit harder to capture once it moved through one material to the other.
But more simulations, more accuracy, more AI engineering. Like this is clearly a direction that things are going. And I think this is just fascinating to hear that people are even working on such a thing. Next up, let's talk about how some new AI research is improving how computers interpret the world. Written by Matt Olsen's. The question here is like, what if the biggest upgrade to AI wasn't just making it smarter like we always talk about, but it was in helping it ignore bad information? Oh my god. Like as a human too, I feel like there's a lesson here for me. I just like I would love to just know right away if information was bad and ignore it. I and you know I don't I've got these heruristics where I think oh that seems fake. I'm going to sort of disregard it or check into it. Or on the opposite sometimes things seem very like yeah I believe that could have happened so I just accept them when they're just not. But researchers at the University of Saskatchewan have developed a new AI method that helps computers recognize actions faster, more accurately and with less computing power. But the problem they're trying to solve is called domain shift. That's when an AI learned something in one setting, but it struggles when the surroundings change.
For example, an AI trained to recognize somebody running in a bright, sunny street might fail to recognize that same action in a dark, rainy park. Their new system is called learnable motion focused tokenization. So, LMFT, and it works like a filter. Instead of paying attention to everything in a video, it removes the background details that don't help and it tries to keep focus on the action itself. And the good news is that means the model has less information to process and as a result it can make decisions faster while also improving accuracy. Kind of reminds me of the Jeepa stuff that um Yan Lun was working on before he left Facebook.
Gosh, I wonder it's been a while since we've heard, but he he raised a bunch of money, right? Let's do a quick check.
Give me a quick update on what's happening with Yan Lun's company in the last three months. Do you think he goes back and like tries to outshine Zuckerberg? All right, so he founded it in Paris. He raised a billion dollars.
Startup spent the last three months laying down its technical blueprint. I mean to his credit, I do like that he's trying something so left field. So yeah, I remember he said Silicon Valley has been LLM pill doubling down against AGI hype. Okay, so AMI actually clarified its commercial targets over the last few months. So what sector it's going after is focusing on areas where physical world comprehension matters most, including industrial process control, automation, wearable devices, robotics, and healthcare. H explain in one paragraph what's different about the way Yan Lun thinks about the future of AI versus somebody who's LLM pill. So Lun's vision centers around world models such as as Jeppa, right? joint embedding predictive architecture which learn by observing the physical world much like a human infant or animal does. This allows AI to inherently understand cause and effect, anticipate consequences and reason through actions before executing on them. Gosh. And then also, what happened to safe super intelligence? Are they like launching anything? Maybe I should do a segment on these videos where I just like talk to AI to figure out what's going on with all the quiet companies. I mean, billions of dollars and geniuses are over there. like where's the updates? You know, the no side quest strategy, you know, to be honest, that feels kind of like what Anthropic did. I mean, they still build forward- facing stuff, but that Claude Code has taken everybody by surprise.
And maybe Safe Super Intelligence is in there doing that. No products, no side quests, no intermediary steps, putting 100% of its compute and brain power towards a straight shot at building a safe, aligned artificial super intelligence. I mean, we I haven't talked about it for a while, but what if, you know, like what if SSI comes out on top? And you know what? That would be like a day where like, oh, Anthropic's not in the lead. It's not Google's race to lose. China did not see this coming, but I don't know. It also compared to the other numbers like SSI's only raised three billion, which obviously is like insane, but also is that enough? Well, they got an investment from Nvidia and Alphabet. Uh, who knows, man. This this whole thing's crazy. Stillia Suscover quietly working out of Silicon Valley.
New hybrid AI model is cutting financial forecasting error across stocks and crypto. So there's a new AI model that beat one of the most common tools for predicting stocks and crypto prices. And it did it in a new way. So this is why it's worth talking about. It combined two different systems or actually what I should say is it combined two different ways of finding patterns each with LLMs that were or not LM but also models that were built to tackle those. So researchers have developed a hybrid AI model CLSTM-HN and it's meant to forecast stocks across indices like stocks and crypto. So most models are pretty good at spotting either short-term price movements or long-term trends, but not both at the same time. So the approach combines two well-known AI techniques. One is called a convolutional neural network CNN. Um, these were actually like around even before the transformers, but they look for small local patterns in data and an LSTM model that remembers information over longer periods, helping it recognize broader trends. Okay, so let me just talk to you about how these are just such different systems. So CNN's excel at finding spatial patterns. They view data like a grid. B like I always thought of them in terms of images, but where a nearby data point is highly related and then they extract local hierarchical features like edges, shapes, and textures regardless of where they appear on the grid. Like taking an image and breaking into a grid. So if you're thinking about stock market, I I really think you got to think of it like a visual. actually imagine the whole ticker, the little mountains that it makes and then think about that as an image and try to find edges, shapes, textures and patterns in that. Now, an LSTM model that views data as an order sequence like text, audio, stock prices over time where the order and context of the data points matter. They remember information from earlier steps to understand the current steps. So, they're much more like, all right, what's the journey that we're on? what like brought us to this location. We have to remember some of the early steps and the order everything went in. That's all important. So yeah, you can kind of imagine how these researchers put those two different models together and then tried to learn a combine a combination of how they both think to get to a final result. And when they tested it on public market data, the model reduced forecasting error by 15 to 20%. And researchers say this matters because financial markets are noisy, volatile, and they suddenly move, making them difficult for traditional forecasting methods to handle. All right, so here's just the straight facts. Police in Australia are now scanning faces live in public in real time. And the biggest debate isn't even the technology. It's just whether the law is ready for it or not. So police in Western Australia have started using marked vans with live facial recognition cameras in public places around Perth. The system compares people's faces in real time against a watch list of about 4,000 people. It includes those with outstanding warrants, registered sex offenders, and missing persons. Now, police say that the goal is to find dangerous people faster while protecting the community, and they say that faces that are not on the watch list are not stored. What's your thoughts on that? Like I I unfortunately can only imagine a world where we have mass surveillance. We just have to hope that the like the community has a really good check and balance system. Um keeping the technology just feels like it's going to be kind of fruitless cuz just AI is going to suck up more things and people are going to invite it into their lives in a lot more ways. But there's serious risks to just scanning everybody through some advanced government system. I mean, not just that the system itself can make mistakes, especially, you know, outside of controlled conditions, but innocent people being stopped or questioned isn't even what I'm most worried about. It's that these systems could just be turned on a dime. And how would we know what if they're not storing it? Or what if that list, you know, you know, it's going to grow to just more and more people over time? And who gets to decide who's on that list? Also, facial recognition doesn't work perfectly on all different groups, raising concerns about, you know, different races or unfair outcomes. Or what if the officers trust the software too much and the system for serious offenders gradually expands for other uses over time. So current Australian privacy law gives police more freedom to collect this kind of information than a private business. And Australia still has no dedicated laws for AI powered public surveillance.
Another reason why this is happening is cuz they just haven't addressed it. The law is not keeping up and it doesn't exist. So people just do it, right?
There's no law against it right now. But it'll be interesting to see because some other countries including the US and China have already made more definitions around this kind of stuff. So we'll see if this is something that is important to the citizens especially like if it's testing civil liberties or you know just take the other side. What if Australia really does become like one of the safest places? They're like look I know there's police vans tracking us all but as long as we're not doing anything illegal it keeps bad people away. It means people who steal or are on these lists are like getting caught more often. Maybe the streets get safer, but you know, it's kind of big brother. It's kind of like fear. Next up, let's talk about the art of literary translation and how it exposes the limits of AI. So, this all kind of comes from a sort of interesting fact that AI is pretty good at translating something very logical, right? Like a legal contract in just seconds. But it can struggle quite a bit with a poem. And if you think about it, a poem is sort of meant to be interpreted a thousand ways, whereas a legal contract is trying to be wording that's very precise. I mean, legal contracts are kind of like the math of words. Poetry is a much more open-ended kind of way of writing. And AI has been not getting good at both of them evenly.
Like AI is getting good at translation, but poetry still exposes one of its biggest weaknesses. And the thing is translating a poem from one language to another is more than swapping words between languages. It's a literary transfer that has to capture some understanding of the culture, the history, the metaphor, the emotional word behind the writing. Probably even have a good understanding of who the actual author is as a person and what they think about the world. You have to step into their shoe, their viewpoint.
And there's some examples in this article where ChatGBT translated the poem and the result sounded smooth, but it missed key meanings for people who really understood it. It simplified rich metaphors. It misunderstood parts of the grammar and it even added details that were never in the original. In one example, it turned a poem that had these layered images of trees in the wording whose quote de becomes drops of silence into lines about glittering diamonds.
like the dew becomes the drops of silence. These are imagery of a tree and it you just translate it into glittering diamonds. It just loses the poem's deeper meaning even if it had a little bit of maybe something there. So even though I think deep in the lat in space AI models really are capturing something human, something much closer to our experience than like next word prediction, but it also does seem like it cannot faithfully carry the emotion and culture like cultural heart of a translator normally would get between poems. So will it get there? I think so eventually. But yeah, as of right now, you can definitely see how it's getting better at sort of needle and hay stack problems. It's getting better at seeing certain patterns, but you know, that abstraction layer still needs a lot of work. And finally, let's talk baseball.
So, AI applied to sports was something I remember kind of early this year and like late last year starting to sort of be a thing. Basketball had implemented some really interesting things. There was an AI that was looking at I think at the Wim Wimbledon, right? like the Wimbledon like with the big tennis matches where it was trying to figure out if a ball like a tennis ball is on the line which is weird to think about but like a tennis ball will like squish when it's hit at that speed and like does even the squish part get over it.
It was trying to be like much better than humans, right? It's trying to make the game more fair. But the strike zone in baseball is wild. Like, you know, Moneyball and and analytics have been such a big thing in baseball and AI has been trying for so long to figure out how to accurately guess or or identify, classify if a ball is a strike or not.
And baseball already has a written rule of what counts. So, you would think, okay, like if we know exactly, you know, what is it? What probably 3T by 3T or whatever the exact rule is has to be x amount of inches off the ground. Why did the Major League Baseball AI team need seven years of testing before bringing it out? And according to Cornell researchers, the problem wasn't just building the technology. It was actually translating a rule that already had to interpret in subtle ways into something that a computer could enforce. So, it's technically called Major League Baseball's automated ball strike system.
So, ABS, it uses a computer vision to detect and track the baseball. It uses multiple Hawkeye cameras to capture the pitch in 3D. By the way, these Hawkeye cameras are high-speed camera arrays.
So, the system deploys a network of 8 to 12 specialized highresolution cameras, and it's meant to capture hyperfast objects, 300 frames per second to be precise. And then they use 2D and 3D triangulation, predictive modeling, but still almost like a poem, the discrepancy, even if it can measure pitches accurately, turning the written rules of baseball into something the AI can enforce, is much more complicated than it appears. Because umpires are constantly making these subtle judgments about how a batter's stance, the posture, the sound of the ball hitting, where the kind of uh catcher's hand is in relation to the swing and the shoulder, all of that stuff actually plays this huge role in when they call a ball or a strike and how the rule is actually applied in live games. So, the technology has limits. That whole ABS system, there's all this like other kind of feeling about what a strike is. Even if the AI was technically correct, they're worried that it wouldn't succeed if players, umpires, fans, and teams felt like it made the game worse, like if they lose the art of baseball. And the weirdest thing is early versions that were tested showed that the strike zone, when it's accurately called, people don't like. And that wasn't something the engineers predicted. So, do you think accuracy is accuracy or do you think we should actually just have humans call it because that's what people like or do you think we need a system that learns all the quirks of the humans so it calls balls and strikes like humans always have AI getting implemented in society? These are the kind of questions we have people. Thanks for watching. Share this video with someone if you think it would be helpful and I will see you in the next video.
Related Videos

Setting up a curved screen with Immersive Calibration Pro 4 and multiple cameras (P3D v4)
FlyerOneZero
23K views•2019-07-21

Robot Learning with Sparsity and Scarcity
allenai
379 views•2025-10-14

Jorge Mendez-Mendez: Unlocking Lifelong Robot Learning With Modularity (2023-10-05)
umassmlfl
237 views•2024-01-06

Northwestern’s MS in Robotics: Student Robotics Projects, 2023
NorthwesternEngineering
1K views•2024-05-31

"Perfect" Turns: Turning by the Gyro - FIRST LEGO League (FLL) SPIKE Prime + EV3 RePlay Programming
ZacharyTrautwein
94K views•2020-10-02

Gorkem Secer: TSLIP-based Deadbeat Running Control of Bipedal Robot ATRIAS
DynamicWalking-wv6qm
298 views•2018-06-22

Self-Driving Cars Need Lessons On Human Drivers | Maddie About Science
skunkbear
26K views•2018-08-21

Milrem Robotics’ THeMIS UGVs used in a live-fire manned-unmanned teaming exercise
MilremRobotics
99K views•2021-05-20
Trending

2.4 BILLION Records Got Leaked...
DeepHumor
15K views•2026-07-22

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Should I buy a Sawmill?
essentialcraftsman
29K views•2026-07-22

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23