DeepMind’s integration of world models marks a pivotal shift from rote imitation to genuine physical reasoning. This approach finally bridges the gap between digital intelligence and the complex, unpredictable reality of the physical world.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
【Geminiの"万能ロボット"は教えなくても働ける】ロボットAIが急速進化「靴ひも」結べる「ダンク」も決める/武器は「世界モデル」→経験しなくても学べる頭脳が完成した【1on1 Tech】
Added:[music] [music] [music] [music] [music] >> Okay, we have Carolina Parada from Google DeepMind's robotics team.
Thank you very much for coming.
>> Thanks for having me.
>> Yeah. So, um I was wondering what what were key milestones Google DeepMind's robotics journey and where does robotics at Google stand today? Could you quickly quickly tell us about that?
>> Yeah. So, we've been working towards bringing intelligence into robotics since the beginning of the team.
So, there's been many milestones. I think just to have a few, we introduced LLMs and VLMs into robotics to help robots understand natural language, to help help them understand their environment, and help them plan and reason. We've also introduced transformers to robotics, which brought this new age of data-driven robotics, and the VLA, which is the foundation model that called vision language action model, enabling robots to understand their environment, understand language, and then transfer the complex concepts into more general actions and behaviors.
Uh we've also introduced what's possible when it comes to learning highly dexterous behaviors, right? Like the example that we put up was a robot tying shoelaces, something that many people thought was impossible in our careers.
And then we also introduced reinforcement learning in simulation for robotics, for robotic control. And we usually work across many different robots, so we showed this on quadrupeds and humanoids and ping-pong robots, and that shows the generalization of the techniques. And then last year we introduced Gemini Robotics. This builds on all of this all of these techniques and combines it with Gemini. And the idea is that brings all of Gemini's multimodal world understanding into robotics by essentially adding actions as a new modality in Gemini. So, you know how Gemini can produce text and can produce code, produce images. It can now also produce robot action sequences, which are essentially trajectories that represent the motion of the robots.
So, those are some of the milestones.
We've also continuous continuously improved Gemini Robotics. So, we also introduced Gemini Robotics 1.5, which brought agents into robots.
That enabled us to go from single instruction following to full agentic systems that are in the physical world that understand the request and then understand how to translate that into solving useful problems.
And we're constantly improving the models.
>> Okay.
And you're basically trying to make general-purpose robots, but what's the point of making robots general-purpose?
>> Yeah, I mean, I think that the opportunity for robotics is massive.
Historically and today, robots are very useful. They're in logistics, they're in manufacturing. We have robots doing really cool uh stage dances and stuff, but I think the opportunity that remains for robotics is to be able to help humans in everyday spaces. And the reason that hasn't happened is because robots are not capable of understanding our world. They're not able to dealing with uncertainty and changes and uh and humans. So, I think in order for us to really fulfill that promise, we need to be able to they need to be able to understand us and their environment and be able to adapt. So, that's where intelligence comes in. If we can solve general-purpose robotics, we could now, for example, bring robots into situations where there is labor shortages, or maybe they could help us in home environments or when there is any other kind of situation where you could amplify humans, right? With with robots that understand them and that can enhance their the the productivity of the space.
Yeah.
>> I think there there there are a lot of the impressive robot tasks like there's some that can do basketball or the the the folding the laundry is all sorts of things. But what kind of tasks are you specifically trying to let robots work on? What What kind of hard work >> Yeah.
>> is trying to solve?
>> So we we definitely when we're solving what we're solving are still fundamental problems about how do you learn and how do you teach a robot to do something that is very general. So we do try all kinds of tasks that are seem toyish, but we also try lots of tasks that are real like these problems of dealing with objects you've seen before. That shows up all over, right? A lot of the the reasons that we don't have more robots around us is because they're not used to handling new objects. So that's an example like the example where the robot is able to do a slam dunk. It shows a few things.
It shows that it can handle a completely new object. It [clears throat] shows that it can handle a completely new task that nobody has ever taught it. So it shows that it's learning from understanding the web and understanding the humans and the concepts of you know, basketball. So it was not about the task, it's about the opportunity that it shows that is this robot can now handle completely new objects, completely new task. You understand semantics because on the doing a slam dunk has a depth to it's to its meaning, right? That goes beyond just exact motion.
So we we definitely evaluate our policies across hundreds, thousands of tasks. In fact, we're limited by the time to evaluate policies as they get more general. So, we do home environments, we do logistics environments, assembly, all of those different tasks.
So, yeah.
>> So, you deploy the robots at home or warehouses or factory already?
>> I would I would call it that we're still in early stages of evaluation. So, mostly we're doing lots of different evaluations in different environments with different objects, but I don't know that I would call it a deployment. To me, a deployment is an experiment. Yeah, definitely. Definitely. Yeah, we like to create many different diverse environments. We never test something only on just one task, one setting.
>> Mhm.
How do you do those experiments? Like the putting the robots at home or >> Yeah. No, we we create a lot of environments in our labs, and then we also work with partners who have their own environments.
And so, together we bring a lot of examples of possible tasks, a lot of examples of possible objects. And and we evaluate the best we can statistical significance on on the generalization capabilities. And we usually for every evaluation, we change everything about it, the objects, the positions of objects. We actually on purpose go and intervene while the evaluation is happening to see that the robot adapts, for example.
So, yeah. And then we always ask people to come and bring their own things, and then ask the robot to do completely new things so that we make sure that we it's not a biased evaluation.
>> Right.
And I was at Google I/O last weekend, and I saw Gemini Omni which improved the multimodal capabilities and world understanding.
So, how does Gemini's progress in those capabilities or agent capabilities translate into robotics?
>> Yeah, I mean, we are always looking at many different research threads to advanced We believe deeply that the field is moving extremely fast, but we still believe there's breakthroughs to be had. So, [snorts] we're always working across many different research threads in order to unlock new levels of intelligence. So, certainly some of the research threads we're working on is what we call scalable learning. What that means is that you can learn without robots. A lot of what you see today is learned through robots being teleoperated by humans. That's bounded by the number of robots you have and the number of hours that you have, extremely inefficient. So, we work on things like world models. So, Omni is a world model.
We work very closely with them and in many ways. So, we try to we have threads leveraging world models for evaluation, for data generation, for general foundation model building.
That's only one thread, though. We also work quite a bit on simulation and being able to push simulation to the point that it can really enable us to do highly dexterous behaviors.
And then we also work on learning directly from humans, from human videos.
Um so, we always have multiple threads going in parallel.
>> Mhm.
>> Yeah. And we one of the advantages that we have at Google DeepMind is that we're working very deeply with these teams.
So, the Omni team and our team work very deeply. The Gemini team and us work very deeply. So, we're we have the capacity to change fundamentally how the model is built so it is works for robotics.
>> Mhm.
What about Waymo?
I think that they're physical AI in terms of physical AI humanoids and robots are physical AI. Also, autonomous cars are physical AI. Are you working with them?
>> So, yeah, we're we've always been I've always been a big fan of Waymo. I think that they're an wonderful example of what of build bringing real robots that have real value to humans and also captures the complexity of doing so. And and a car and self-driving car it still has a lot more structure. It still has lanes and traffic lights and rules of the road. When you go into full humanoids, there is no laws that define how you should move around the space, right? So, it's a lot more complex. We work with them in a few areas. One of the areas that that we're learning a lot and we're really excited to have them as a partner is on safety. So, they've always put safety first in their strategy and actually has become one of their um leading capabilities and leading reasons to use a Waymo, which is excellent and we want to do the same. So, we also have a very safety-first approach when it comes to developing these models and so we're learning a lot from Waymo.
But yeah, generally we also in interact when it comes to world models. They also were using world models only for for doing a lot of I don't know if you saw their announcement is incredible. It's a great example of why you want to learn some stuff without having to experience it. You definitely don't want to put a Waymo in front of a tornado in order to understand how the robot would behave would behave, right? So, you can generate that with world models.
>> Okay. And what role do you think humanoids will play in the future of robotics? And how does Google DeepMind your robotics team work on humanoids?
>> Yeah, I mean, I think I mentioned this many times is I believe the world will look like a pretty large ecosystem of many different robot types, not just humanoids. There is some tasks for which other robots are actually better suited.
So, we want to build the intelligence that works across many different robot types.
Having said that, I do believe humanoids is going to be an important category within the set of robots in that ecosystem.
I find them super interesting for the traditional answer for why humanoids is well, they can go in any human space.
The world was designed for humans.
Humanoids can just walk into any space and fit in. And that is completely true.
It removes all the boundary all the constraints that many people have to deploy robots, which is you have to modify the environment. With humanoids, technically, you shouldn't. I think there's a more interesting reason to use humanoids, which is humanoids reflect our human form. And when you're learning and you're trying to learn beyond just moving robots, humanoids are so close to humans that give you that opening that all of a sudden you could just learn from humans how to how the humanoids should behave. So, that's one point that I really find fascinating. Another reason why I like having 10-finger hands, human have 10-finger hands. You can learn from humans how should you do different tasks just watching them, right?
So, that's one reason. The other reason is that we are very keen to go all the way towards AGI in the physical world.
How do you show AGI? I think at best a great example is to show a humanoid doing everything that an average human can do.
That is indisputably a a good example of you have reached AGI in the physical world. So, that's another reason that I find it interesting to work on humanoids.
>> And when are we actually going to see the robots deployed in everyday life?
>> Yeah, I get that question a lot. I think you will get different answers from different people.
My view is that the intelligence layer is going to move very very fast.
I would say definitely to solve fully robots in our homes, human robots in your home, it would take a long time.
Like past AGI because when people talk about AGI, they're talking about AGI in the digital world. So, in the physical world, I think it's another step. So, definitely past 5 years. I do think that the capabilities will evolve quite quickly. We're seeing So, a robot that understands the environment and is able to to task and adapt and all of that, I think it will happen, yeah, not too far beyond the 5 years. The problem is going to be are you also able to run this robot reliably, safely, with privacy, security, and safe you know, all the levels of safety, and all of the approvals to run in a home, that's a different story. So, my guess is that it will happen at about 10 years. So, that's just my guess right now. But, there's always a range. Things are always moving faster than than we expect, so.
>> Yeah.
>> Yeah.
>> I have one last question. Since you're here in Japan, uh but one of the NVIDIA's robotics team leader uh said that Japan is no longer a robotics superpower. Well, I would I was wondering what's your take on that, and uh how do you see the opportunities of research or collaboration opportunities here in Japan?
>> Yeah, I mean, Japan has a really long history in robotics. I mean, is I would say the Japanese is a superpower in robotics, right? Like some the major companies that have the most robots deployed are in Japan. Are from Japan.
So, um I find that I'm super excited to engage more with uh with the Japan robotics community. That's why I've been here for a couple of weeks talking to different companies. Uh we just uh engaged with many many Japanese companies for a trusted tester program.
We're working with with Fanuc, with Yaskawa, uh KDDI, with Toyota. So, we're really excited. It's early days, uh but they're definitely very engaged, and they're very interested in bringing this new wave of physical AI um and combine it with all of their expertise, right? I think a lot of what's happening in robotics today is happening on prototype robots that are not don't quite have the maturity. I think the superpower for Japan is that they have this experience of bringing really reliable robots that are very precise and that work for long periods of time. I think that's going to matter when we switch from just prototype and pilots into real production. So, I still think that yeah, I don't have that same opinion.
Definitely don't share the same opinion.
>> What are you doing with those trusted testers right now?
>> Yeah, so our trusted testers is a way for us to get to start working with a company. Uh we give them access to our models so they can evaluate early our models and then they can tell us how it works. They also often times really engage more deeply and bring us benchmarks so that we understand uh the harder task that they care about. So it's a great way for us to start get gathering feedback in a way that engages us with the company. And most of our deep partners started as trusted testers and then evolve into uh stronger partnerships.
>> Mhm. Okay.
>> Yeah.
>> Great.
>> Thank you for having >> Tell me. Thank you.
>> Thank you for having me. Yeah.
>> [clears throat] >> Thank you very much. Thank you very much for your time. Thank you. Yeah.
You're here in a couple weeks.
>> Yes, a couple weeks. So this is my second week. So that's why I'm adjusted today.
>> [laughter] >> It took me 10 days.
>> lag anymore. Okay. Yeah. Yeah. Yeah.
Related Videos

Setting up a curved screen with Immersive Calibration Pro 4 and multiple cameras (P3D v4)
FlyerOneZero
23K views•2019-07-21

Robot Learning with Sparsity and Scarcity
allenai
379 views•2025-10-14

Jorge Mendez-Mendez: Unlocking Lifelong Robot Learning With Modularity (2023-10-05)
umassmlfl
237 views•2024-01-06

Northwestern’s MS in Robotics: Student Robotics Projects, 2023
NorthwesternEngineering
1K views•2024-05-31

"Perfect" Turns: Turning by the Gyro - FIRST LEGO League (FLL) SPIKE Prime + EV3 RePlay Programming
ZacharyTrautwein
94K views•2020-10-02

Gorkem Secer: TSLIP-based Deadbeat Running Control of Bipedal Robot ATRIAS
DynamicWalking-wv6qm
298 views•2018-06-22

Self-Driving Cars Need Lessons On Human Drivers | Maddie About Science
skunkbear
26K views•2018-08-21

Milrem Robotics’ THeMIS UGVs used in a live-fire manned-unmanned teaming exercise
MilremRobotics
99K views•2021-05-20
Trending

we're almost finished the house (ep.125)
JennaPhipps
347K views•2026-07-22

We Finally Know Where Saturn’s Rings Came From
astrumspace
79K views•2026-07-22

BIG BET: Cathie Wood goes ALL IN on Elon Musk
FoxBusiness
89K views•2026-07-22

MIC DROP: Smithsonian Director Called Out For Woke Propaganda
TheAmalaEkpunobi
37K views•2026-07-23