The talk effectively exposes how hardware constraints and latency, rather than just intelligence, are the true gatekeepers of the VLM versus world model debate. It’s a necessary reality check for an industry often more obsessed with parameter counts than real-time physical execution.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
World models, VLAs: performance, accuracy and perspective
Added:Hi everyone. Today I want to discuss a modern robotics and AI in robotics with Gleb who is a founder of robotic learning collective which is like a community about VLM and world models and I have a lot of question to him about which model is the best how we can compare them and a lot a lot of interesting question uh if you're interested in robotics uh definitely this interview can give you reference point where you need to start your adventure and what is the most interesting right now. So let's start.
Hi Gl uh thank you for coming and uh probably can you tell in a few words about your uh research lab? What are you doing events that your organization?
>> Sure.
>> Sure. Um so we are called Robert Learning Collective. uh the the very loose association that I co-ounded a year ago and we mostly focus on just doing some cool projects in robotics in the space and sharing knowledge learning something ourselves and um we'll be we have been working on two big projects one of them was um the work on the simulation um on the models trying to solve domestic task in simulation. So it was uh last year and was probably one of the biggest competition in BLA space and uh it was part organized by Stanford and was part of Europe's and uh we took the first place by applying modified version of the physical intelligence model to solve multi- multitask scenario where a robot need to go around the flat complete several task task and score uh how many task the robot completed. So that was one our big project. The other big project that we worked was uh trying to make VA models smaller. So again last year we saw some interesting result that um VA models can be used in a more natural way uh in a more natural way use VLM model itself. So you take just VLM model which is basically standard language model plus visual encoder and unroll it out regressively and we show that um such models can be significantly reduced in sized and basically um reproduced one of the paper in build a paper but instead of seven billions train model of the half billion parameters and obviously that's makes VA much more accessible to everybody for experiments and um this year we are focused on um running regular paper clubs. So um I don't know what kind of different people have different experience but I was privileged usually work in very um high density talent environment and there was always some people who were very passionate about learning new stuff and somebody organized paper club where you meet I don't know once a week once every two week and you discuss new research and I run such uh sessions myself in my uh previous job and when I when I left it and have time in between uh we decided to organize the similar activity but open to everybody to come in London and uh we tried also to make up online uh translation and this is a bit unique because there is not many discussion paper clubs especially in robotic space um and uh so Far we've been running like for half a year almost like quite successfully in term of what kind of speakers we have uh have speakers from AA from humanite from wave um and in term of like how active our discussion is so I'm uh very very happy that we managed to do this.
>> Yeah I also joined I think four of your clubs. It was super interesting because uh this field is super big and it's really hard to understand like in general what is happening and it's super nice. I think it's unique event. So this is why I decided to organize um this interview because it it was really great and I found for myself a lot of new ideas how everything is organized. So maybe one of uh so our first topic uh will be about uh this and uh I think your four initial event were more focused on uh world models and uh let's discuss a little bit what is the world models because everyone is uh asking in what is the different from VA how they can help are they really help and so on so what do you think about them like your impression Maybe I understand that it's possible to speak infinite about this but maybe like some short idea.
>> Sure. Sure. Uh I want indeed we run like four session on this and it's forced me to cons to think harder about this research more and uh at the end consolidate a little bit my thinking. So um I probably can provide at least some uh some thoughts on this but obviously I'm not a like an expert. Um so in in a very simple terms uh VA's model is the way to get the best from the LM world um in bringing to robotics.
So you kind of the whole idea of having VA primarily came from Google when they were working on RT series models and it was a little bit of strange idea to take like VLM model which is basically a language model that also embed images in the language space and try to teach it to do robotic stuff which was completely different from what people used to do in robotic space because robotic space was all historically about control. Uh uh so it was even pre-model era of um model based control. There was also a lot of interest in reinforcement learning and a lot of reinforcement learning came from robotics and was uh was developed first there. And these models were never pre-trained on something like language or images. And they were usually very small so to run them because you want to run them fast etc. And anyway you don't have a lot of data so why do you have why you want a big model so the whole idea of using VLM was very strange but it's actually turned out to be quite quite reasonable uh idea and quite successful right now.
Obviously it doesn't solve anything everything but it became kind of let's say the paradigm right now in robotics when we talk about manipulation so va but at the heart of it you still have this language model plus visual encoder and word model came from completely different world probably a bit more connected to origin of the robotics or origin of the control theory array that there is a world outside and you can model it and if you can model world outside you can basically uh unroll different varants varants of how you would act and basically do planning. So if you if you remember something like uh Alph Go and museum papers they also have some kind of way to predict how work would change.
But in that case work was either uh spec have specified property like chessboard or goboard or some kind of abstract but learned from this quite simple representation because chessboard and go board they still very a very limited number of combinations. No combination is enough but limited number of uh states.
So this idea of world model of that we want to model reality environment around us and want to plan in in this reality or maybe even learn how to how we should act in this reality uh was was quite also as I said like coming from control theory more and uh the way people do it right now they take a video model so the model that understands and video and can produce video as well and people apply various technique how to modify such models. So you can actually control them not only through text which gives you some control but it's more like you just prompting all dream me a video of something and something but instead people modify this video models to make them reactive to uh high frequency changes like agents going left turning around moving arms etc. And this way you get one of one form of video model. And yeah, we see last year probably even less last half a year a lot of interest in this in both models and they kind of start to challenge dominance of VA model and some people already say that VA models are dead etc. But obviously sedation is much more complicated. There is much more in between state of these two world. Um but yeah.
Yeah. uh and I agree that situation is super complicated and maybe one of my first question uh to world models it's like they're pretty big and for example uh samples like articles that was on your paper club usually it's like you need at least one 40 GB uh GPU just to run this model uh what do you think is it like uh the real problem uh with them or will they be a a little bit smaller because like what we see with general VAS like NVMs they became smaller and smaller faster and faster so do you think it's possible for world bundles because maybe it's one of the main limitation for them >> yeah it's a good question because in a sense work model right now is the easiest way to trade compute for u data so we If you have less robotics data, especially less data from your domain, you can basically instead of VA or let's say behavioral model, you can use word model and it will be more expensive to train, more inspective to run, but it somehow uh require less data. So that makes sense.
Um but I'm I'm quite optimistic about the scale obviously right now. Yes, it's it's a bit unaccessible but even right now there are it's a range of the model. So for example, Google deep mind in the dream dreamer 4 which is very good paper uh used two billion model uh two billion parameters model which is somewhat okay some in the range of typical VA that we have right now anyway. Um there are bigger model that use 14 billions like dream zero paper from Nvidia for example and yes it's much harder to it's super expensive to train them and even much harder to run them in real time although it's possible as they show um so I would say they probably it wouldn't grow in two direction there would be both like we would just see the more diversity the same situation that we see with VA there will be smaller model but in bigger model as well uh as the technology evolve as the people find more and more uses for this you will probably have much smaller world models although they still should be bigger big enough to feed the whole environment right the whole dynamics of the environment but I I I I I would expect to them to become a bit more accessible but also the other end is completely uh unlimited how big work model can be and we know that the companies experiment with quite quite significant work models.
Yeah, I I also saw maybe a week ago uh Li robot appended uh Vij Japa 2 and uh I was like a bit surprised that it's pretty small actually it's just uh ba based at least inference part like they have I think bigger uh training part but with inference part it's like the regular small like 2B v uh VLM model and I don't know if we can call this world model. Of course, it's like maybe a little bit tricky, but it's pretty interesting here.
>> Yeah. And compute is we're getting more and more compute and it's becoming uh in relative terms cheaper. Like the first VA that emerg went like developed two three years ago, people say that oh it's impossible to run for anybody unless you're Google. And right now a lot of people just at home as a hobby project run some kind of BLA.
>> Yeah. Yeah. Um maybe even like the robot a little bit democ democrazed all this stuff because like it became like you just need to run one script and you can train a model. It's super easy. Uh what do you think about like maybe accuracy of world models because like yes there are some articles where we can see some benefits in accuracy usually they are not super big but as far as I understand the m main like promise of all these world models it's like they given us uh ability to work in new environment more stable to new data and uh so in general are there any stable metrics that can show this?
Uh do we have a really big gap from VAS right now in terms like of accuracy of general ability?
What are your thoughts?
>> Yeah, unfortunately situation with metrics and benchmark is terrible at robotics uh because there is no um no shared benchmarks. There there are a couple of like benchmarks in simulation but everybody hates them. everybody think that they are very bad. Um so yeah it's it's quite shame that we don't run some kind of competition where the robots are the same or maybe each company can bring its own robot and show what they can do and basically we are reduced to judge the progress by either demo videos which are completely cherry picked or by maybe the best way maybe the best that we have is like some technical reports from companies that usually also compare different things with different things. So it's I was trying to find proper comparison between BLA and both models and it it was very hard to to find because again you want probably something like the best effort in BLA versus the best effort at work models.
But the team that is building this paper they would be either doing one or the other and they therefore would under represent the other side. Um we did find like ultimately find some not direct comparison.
Again, you can you can check in in in Dreamer 4, although it was in simulation and even like on Minecraft, but I somehow still um find it's pretty pretty convincing how they approach comparing and um in term of metrics again. So there is it's it's a bit hard because there is metrics for the work model itself how good it predict the world.
Uh but so how good it predicts the world and the second metric how how good your robots uh succeed in the in the task that is in front of them. And obviously if we talk about tasks it's immediately become uh very complicated because it depends on what task in what environment how much data it was trained for this particular task etc etc. So although there is pretty understandable metric of just success rate for this task, it's it's impossible to compare but we do see several points that indicates that yes world models in general are better generalized. So you need less data even like zero short.
Um, regarding the the more uh strict metric of how world models evaluated themselves um if if you enroll them and try to predict how the future would look like the there are kind of several um options. One of them is just pixel difference. So right if you predict the future and this is this is somehow how hold out example you can just compare pixels but obviously it's pretty bad because the future is multimodel you can just throw slightly differently and it's still the same room but pixels are completely different so people use some other metrics I always forget the name but there is this clever way to compare not raw pixels but feature space in uh in um vision encoder. So you just take a visual visual model and surprisingly if you make this metric it's very much aligned with what human think are similar pictures.
Um so people have laid board models like this but it it's quite tricky because yes of predicting future is inherently multimodel so you can predict two different pictures and both can be simultaneously correct and there are ways how to work with it. So how to construct metrics that take this into account. So you know several row times and your ground truth should uh align with at least one of your prediction right so let's say best from K um prediction or something like this but yeah from the research that we see right now from the companies and from academia especially from the companies who have money to do this we have nothing nothing too much of the rigorous like at least that they share. So it's it's it's pretty like we done this, it's working for us. Here's some numbers. You cannot reproduce them.
You just have to believe.
>> Yeah. Yeah. Yeah. Uh another like maybe issue with them when I'm looking on some articles uh I have a feeling that uh in general they are a little bit overengineered right now because when I see like a thousand of different part not the thousand of course like dozen different part of the same model of the same process it's always a question for me how good it was optimized for the specific task uh because it's like not end to end uh process And uh in this term VAS they maybe a little bit more like uh um simplified and easily like I can understand what happening inside like VA much easier than in world models. So it's at least for me it's always a question is it like overgineered or not >> and maybe what do you think?
Yeah, the the papers quite often overengineered because yes, I for example I I usually have some simple question like is the world model better than VA? But usually the authors they try to answer this plus some other questions. they have some clever ideas how to do something I don't know planning or similar or training in the work model or something else or some new way some just clever trick to do something and they include this as well and you're like why is it needed is it just because they have this nice idea cool idea and they wanted to try yeah it's it's usually a mess but that that's that's why work with resources is is such complicated to to extract all this.
>> Yeah. Yeah. For me it was super funny article uh last year when I saw like uh uh VA zero articles where they just took quen and uh predict like uh the same and found that it's working better than VA and it was like super funny in this term like how of course like latency and maybe it's a good like next question about all this latency issues. Uh do you think they are solvable and at least at this stage? Of course if in future we will have better hardware maybe with it we can solve but right now we have a significant uh latency uh from the model uh side we have significant latency from hardware side because even like uh if you're looking on how fast camera can deliver images to Jetson it's usually like around 50 millisecond yes it's possible to do faster but like regular it's 50 millisecond and uh what do you think about all this latency issues maybe not only related to world models maybe to the LA as well and how good it can be solvable >> yeah that's a good question so yes it's obviously very uh painful to work to work with system that have such high such have high latency And for really excellent performance you need system to have be very reactive. Um but I would say this is quite an engineering problem and u engineering plus time and money. uh even more less time less money more time in that sense that for example in word models yes they are big and slow to run etc because you need to predict images but there is a lot of tricks that you can do some of them is just general approaches to speed up any model inference and some of them uh more uh related to word models so for example in the dream zero from Nvidia they took 14 billions model which is significant model and uh I think they they make it faster like 40 times or something like this so they may make it to be able to run with 150 millisecond latency which is which is insane it's super fast they started with six six second or something like this and get under 200 millisecond. Uh and they use like classical approach like general general approaches like uh compile with um um CUDA graphs uh lower quantization uh just bigger and faster GPU because they have uh they run they run this model on the two uh black GPUs B200 if I'm not mistaken. So if if you really want to uh squeeze squeeze the latency, you can do this.
And if we talk about word models, uh yes, it's very expensive to produce images. But you actually you I say produce images, you usually don't even need to produce images. You just need to produce latent uh latent space of these images. So it's already smaller and also there is a number of work like from mimic and from Nvidia as well.
Yes. Uh that shows that you don't need to produce fully denoised images because you can so how images produce it's it's a diffusion of flow matching process. So you start from noise complete noise and iteratively applying your model you get to a very sharp realistic picture right so in that sense you don't need to finish the whole this process before you actually start to run more like model submodel or components that are from this image try to predict what action should be and this is significantly again can like improve latency I don't know 10 10x or something like this.
Another another example is again dreamer 4 from Google Google defined where they also speed it up like 30 times and then make it bigger. So it's total speed up would be probably more than 50 times and they comfortably run it with uh real time on one H100 which is expensive GPU uh not already not the latest generation expensive GPU so we probably don't have it at home but it's still like if you're a researcher you probably have access to it so um in the sense performance is to reduce latency is is always complicated because you basically fight against entropy. uh you solve kind of biggest problem and then you see that oh there is like bottlenecks on other layers etc etc and but on the other hand reducing working on something like performance and latency is very easy because you have target you have something like you don't need to oh what should be what should be our metrics how how we will envision our solution you just have something working and you just try to make it compatible with quality but faster. In that sense, we know how clo is or any AI system good at if you just give them something measurable, they quite good at optimizing it and humans are as well.
Yeah, I agree in this sense like uh that in future it will be much better I hope to optimize uh and uh also a lot of hardware I I hop in that Qualcomm will do the chips that will be like will be able to run VAS maybe other vendors as well and uh because right now it's sometimes it's a little bit sad to see on all new project where it's like a pretty expensive robot which is definitely connected to one expensive GPU nearby uh and it's making things like definitely unbearable for most of application. Yes, for research definitely it's okay but when we're speaking about some real application real world it's a little bit uh more questionable. Uh by by the way uh did you have any maybe experience with such optimization for hardware or uh you do all experiments and uh stuff on Nvidia GPUs?
>> Yeah, mostly on Nvidia GPUs >> and the thing is that yes you can you can improve it. You can always like you can get made custom custom hardware like the thing that Tesla have been doing for quite a while. So they don't run the um models inside the car on Nvidia GPU.
They h they they developed hardware themsel and it's much more optimized in terms of like power in term of uh cost etc. But with all these problems, I think in robotics we yeah it's still like we are quite far away from this thing just working. So it doesn't matter that much to optimize to optimize the for for for I don't know compact size or something like this maybe soon. But again, if you will have something working right now, you can always go to VC and say, "Hey guys, I have something that is working.
I just need to make it smaller. Give me money." And if you can show that it's already working, you will get money and make it small.
>> Yeah, I agree. May maybe expect uh accept uh gumoists because for gumoids it's pretty often important to like uh when you have a low uh high latency on your Wi-Fi it's like creating additional loop here and maybe for mono it's better to have something local and >> yeah you definitely need something local like let's say model that control the commotion and stuff like this.
But logo stuff can be relatively small like like locomotion models are usually quite small.
>> Yeah. Yeah, I agree. Uh what about all this like VAS, world models and uh other stuff uh with a different data than uh just row images because in my opinion it's uh an approach at least where hardware robot technique is moving last year a lot there was a lot of uh new sensors uh some touch sensors I near ultrasound sensors uh how do you think it's it will affect all this uh because we just can't generate all this information with world models is it like uh a limitation and for such data only like classical VLMs are available and do you have some thoughts about this >> yeah that's that's a big question there are several way how to approach it So yeah, let's let's just paint the background. So right now all VA models and board models they consume just video uh which is images.
So everything run on the camera. Why it's so because cameras been used extensively even before they start to be used for AI. Uh so they become very very cheap, very reliable, uh very standardized. So when you run your model, you don't ask like oh what camera produc this like image you just like oh it's it's it's an image I I can work with it because we have pretty standard format for it etc. And um obviously for a lot of people who are coming more from like classical robotics perspective this is very strange that we use only cameras right now because there's so much information in the tactile sensors in the um force sensors etc. uh but um the the picture is I would say layered layered. So for example there is uh quite famous video experiment when somebody do experiments on human it's and they anthize basically freeze the the sensitivity on the hand and ask these people to uh to use matches right to strike the fire and people struggle a lot because they they don't feel when they actually hold the match etc. And uh the time it took to light the match is like on the order of magnitude or more than if they have their properly working hands. So like this kind of tactile feedback feedback is definitely important for human and it's definitely important if you want uh very low latency very robust models that for example when sometimes people like I I can go into my pocket and find something I I don't I don't put my my head inside my pocket to find my phone right um but so why I'm saying this Because I I do agree with this and I do believe that in the future when we have let's say humanoid robots doing everything that human can do at home or something like this they would have sense of touch and for sensor etc. It's all very needed but right now we have problems even more uh more fundamental. For example, when you see this experiment with on people with frozen hands trying to strike a match, they still succeed and they also understand what need to be done. They just struggle with the with like with the high precision manipulation, but they don't I don't know try to do some some other task with this much.
And um current models they robotics model they still kind of live in the age of probably CH GP2 GP3 at the best and uh it means that sometimes they just do something completely something different so they let's say hallucinate or whatever and in that sense even human without tactile sensor is is much better right now.
Uh somebody may say that there is like a bit more complicated argument that if you have this all these sensors the representation that you you would learn would be much more robust and better and this will allow you to actually um learn better generalize and not not do some random stuff. So there is argument but I I I don't fully fully believe in it in the sense how how it's solved right now to certain extent you have cameras on the wrist of robot which is very close to the hands and it's kind of substitute for tactile sensor and also you can use some material that on the end endector that is slightly bendable. like deformable. So if you actually touch something, you can see on the camera that oh I'm applying some force and model theoretically you should be able to learn this. Um so I would say I'm a very practical person in that sense. Let's let's make let's say human can teleate robot right now without sense of touch and they would struggle with some complex um fine detail motions but let's let's get at least that level of performance from the model and then go further in that sense probably the best example that I personally experienced was that I've been working on self-driving quite a lot and in self-driving If you get a bunch of engineers in the room and say, "Guys, I need you to devices a self-driving car, they would start immediately going to like, okay, what what can we should what should we safeguard against the the smoke, uh the the the rain, the um light shining, the total darkness, etc., etc." and they would devise just insane amount of sensors that you need to put on the car. And if you look at the I don't I haven't checked like last year way car but if you look like three years ago way carar is just it's it's surrounded by sensor I don't know they have 12 or 14 lighters um because for every place that can be oluded they put a lighter there and probably two to overlap and and have a better picture and you have this approach but you have approach of the company that where I uh worked which is a UK company called uh Wave and the first iteration that we done was with only one frontal camera. We even have other cameras but we have some models working only with frontal camera and you would say okay you have just camera one camera looking at front and this model even doesn't have memory. So we just like okay what I see inside in front of me I will try to drive and uh this shouldn't work but it was working in the sense that it was able to drive like for a substantial amount of time without human intervening and you you will be able to show to investors that guys we already solve the problem 98%.
We just need know a lot of money to solve it 100%. And that's what they done and quite successful. So in that sense for me it makes total sense to solve the big big chunk of the problem first and and then add more sensors if they need it when they would be cheaper in the future etc etc. Yeah, I agree that uh different ways to solve the same task and um each one of them like maybe faster for some conditions. So initially maybe Tesla was a little bit faster in terms of how they develop algorithms but nowadays maybe like position of way that you mentioned is a little bit more stable in terms how general they can be so and we don't know how it will change in the future. What like maybe the final question, what are your prediction about when we will have like uh all these robots with VAS or world models uh in the real life because nowadays if we look on some factories, it's usually even more classical robots and more classical uh models and if we look on some uh humanoid, usually you just can't buy like I think only one X robotics are selling like fully capable robots that you can just put in your home and it will do at least something.
So what is your expectation about progress like is the current level of technology enough to make at least something useful to you?
>> Yeah. Uh many predictions are very very uh dangerous. Um so don't trust anything that I would say but from the from the some work that we have seen recently for example physical intelligence they show some uh work that they done with collaboration with more like deployment on some kind of factory not factory but places where they package package your package Uh basically wrap your package in some kind of bag and um surprisingly like it's 2020 2026 and uh human haven't solved such such problem uh yet um at least in in a general sense. So you need AI to do this. Um so from this work and from just amount of money that they get from investors uh other companies as well like generalist and other like Dina Robotics, Genesis will probably see some limited commercial pilots very soon. I I would expect either I would definitely expect them in a year.
So some commercial pilots where there is like couple of robots stationary and they doing something sorting packages or whatever because if they wouldn't do this I would I would um they would probably have too many questions from investors but it would still be quite limited in the sense that when started they also started in one city very limited hours only when good weather only on these roads etc etc. So yeah the the the easiest way to predict is just to to say the future would repeat the past and say we will see something like uh self drive maybe it would be a bit bit faster because the safety concerns are lower that's one of the reason why I moved from self to robotics because I don't want to be slowed down because my robot at least I I can I can um test it whatever I want without being afraid to kill somebody. Um and also so because the safety problem is is is less extent at least you can run robots on factories and warehouses and also because you know we have cloud or whatever you have right now. So development is a hopefully a bit faster and also there is quite a lot of interest from VC and just amount of capital. So maybe maybe we will get let's say for self-driving it took from the start 20 years right or something like this till at least you have more than like a handful of cities on the world where in the world where they are working. So hopefully we will get something faster in let's say in 10 years. We will see it will be okay to have some advanced warehouses or factories using using humanoid robots.
Maybe not humanoid in the sense that it be looking like humanoid but you you get the idea like mobile with hands etc. >> Thank you Ple for your uh talk and it was super interesting to listen. Thank you. Yeah, thanks so
Related Videos

Setting up a curved screen with Immersive Calibration Pro 4 and multiple cameras (P3D v4)
FlyerOneZero
23K views•2019-07-21

Robot Learning with Sparsity and Scarcity
allenai
379 views•2025-10-14

Jorge Mendez-Mendez: Unlocking Lifelong Robot Learning With Modularity (2023-10-05)
umassmlfl
237 views•2024-01-06

Northwestern’s MS in Robotics: Student Robotics Projects, 2023
NorthwesternEngineering
1K views•2024-05-31

"Perfect" Turns: Turning by the Gyro - FIRST LEGO League (FLL) SPIKE Prime + EV3 RePlay Programming
ZacharyTrautwein
94K views•2020-10-02

Gorkem Secer: TSLIP-based Deadbeat Running Control of Bipedal Robot ATRIAS
DynamicWalking-wv6qm
298 views•2018-06-22

Self-Driving Cars Need Lessons On Human Drivers | Maddie About Science
skunkbear
26K views•2018-08-21

Milrem Robotics’ THeMIS UGVs used in a live-fire manned-unmanned teaming exercise
MilremRobotics
99K views•2021-05-20
Trending

we're almost finished the house (ep.125)
JennaPhipps
347K views•2026-07-22

We Finally Know Where Saturn’s Rings Came From
astrumspace
79K views•2026-07-22

BIG BET: Cathie Wood goes ALL IN on Elon Musk
FoxBusiness
89K views•2026-07-22

MIC DROP: Smithsonian Director Called Out For Woke Propaganda
TheAmalaEkpunobi
37K views•2026-07-23