Contact grounded policy is a robot learning approach that uses three coupled variables—robot state, tactile sensing, and control target—to bridge the gap between high-level predictions and low-level execution, enabling robots to perform dextrous contact-rich manipulation tasks like jar opening, box flipping, and dish wiping with approximately 15% performance improvement over baseline methods.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Ep: Contact Grounded Policy
Added:Hey everyone and welcome to another episode of Roa Papers. Today we're excited to have Yeping and Chong Tong to talk to us about contact grounded policy uh just visual textual policy with generative contact grounding. We'll learn about All right guys, could you tell us could you start off by telling us a little about yourselves?
>> Okay. Uh maybe I can go ahead. Uh hi everyone. I'm John Kong. I'm a PhD student at Purdue University. Uh I studied uh robot learning specifically.
I'm interesting. I'm interested in how can we leverage multi sensory data like a vision plus touch to learn um text recognition policy. Uh yeah, that's pretty much my my research and uh myself.
>> Yep. Uh so I'm Yeping and I'm a PhD candidate at University of Wisconsin Madison and my research is mainly about robot manipulation including learning methods and the traditional motion planning methods. Yeah. And uh uh this work was done uh when Junto and I did an internship at MATA.
>> Awesome. All right. Well, let's take a break. Could you tell us a little bit about what you worked on?
>> Yeah. So, uh very excited to be here. Uh and uh this is our work contact grounded policy uh dex tactile policy with generative contact grounding. Um yeah so this is a teaser of the project.
Basically uh we can uh I mean uh our our method can enable a robot hand do this kind of like opening task which is uh contact ridge you can the robot need to leverage the mot finger to be in touch with the jar lid and try to open it. So uh yeah I guess uh um yeah this is a so I want to talk about the motivation of our project. So let's start from this diagram. This is a very common structure of many robot policies. So the robot takes observations such as images. So images is most common but uh some policy also take um you know touch signals as well as the robot states and the policy sorry and the policy predicts actions which are some kinematic targets for the controllers such as the post or joint angles but the one important part is also often so robots do not directly interact with the world through positions uh robot actually interact through the forces and context. So if we only predict targets we may miss some of the k physical interaction that actually determine the behavior. So let's look at two examples. In the first one we train diffusion policy for uh and uh in this task in different rows we actually randomize the oh okay we actually randomize the table height to add the difficulty. So we observe that the policy often applies uh too much force on the dash instead of adjusting the wiping behavior based on the table height if we only have the uh kinematic target as the objectives for learning.
And uh yeah so the pop robot kind of get stuck here. Uh and in the second example we use a diffusion policy for box flipping. So this is um like nonprehensial management task. Uh so here robot need to leverage the the whole hand to try to flip this box. But uh uh and we can see in the video if we only if we only have the kinematic prediction um sometimes the robot cannot apply enough forces for flipping this video.
So uh we think this is the gap between the uh kinematic prediction and the actual contact reasoning. So uh in our projects so our goal is try to find a method can uh make the robot do all this kind of like uh dextrous contact rich magnetic task uh including um like the jar opening I showed as well as the box flipping and also like some some dish wiping some uh fragile object grasping tasks that's the uh motivation of our method. Uh yeah. So >> sorry. I have a question. Uh is the the dish wiping is the simulation. What simulation is that?
>> Like I'm just curious. Pretty cool. Like you're simulating the the plate being dirty.
>> That's a good question.
>> Yeah, that's a very good question. So uh sorry, I don't know why this is playing autonomously. Oh, I guess because of this. Okay. Yeah, let's Yeah, the the this is a very good question. So this is actually not uh like common simulator like Mojo or is lab. This is actually Unreal Engine which is a game engine. Um but it actually runs some uh physics physics engine back end um developed by Meta. Uh so it's uh yeah you can think about this is a in-house simulator.
>> Okay.
>> But uh yeah but I mean um yeah yeah it's not open open yet. Uh but I think it has some cool features because we can see here I think because of the UEIE you can simulate a lot of stuffs. For example, you can actually wipe this dish. Um um because basically in the in the video game you can design everything right. Uh and also uh the back the physics back end also provide us some some of the soft body simulation. If you take a look of this bond, this yellow part is soft.
And also um in this video you can say this actually breakable. So if you apply too much force on it, it will break. Uh so I think all this like >> if you turn over the sponge will be dirty.
>> Uh put over shaking his head.
>> No no like let's say for example you see the the plate is dirty, right? So you if a sponge is clean, you wipe it. Would the sponge be nice? So uh yeah this is more about the details about simulator but I think uh this is actually not >> yeah [laughter] >> I think is it actually printing so so I think the way to achieve this simulation is actually printing like white color on this plate >> I was just distracted by the simulation >> yeah I think I think it's is a very very fun simulator so and also because yeah because of all these features it actually allow us to do a lot different evaluation in simulation because I think sometimes in real world uh it's very hard to do the large scale evaluation uh and also you know the contact regulation is also always a challenge for evaluate um but I think because of this kind of breakable egg as well as this kind of awesome soft body simulator as well as this kind of like dish wiping environment we can actually evaluate everything um like in a large scale so which is good okay uh yeah so so next I want to I want to So um yeah some of the fundamental like uh uh um fundamental uh theory of our method. So we actually start from a very simple idea in complant control. Um I show in this diagram this one. So you can say uh this is a very common behavior when we have the uh compl compliant controller uh and the robot being contacted with some object. So basically when the robot touches the object the controller target and the sorry I don't know why it's always autonomous play but yeah sorry about that uh yeah let's let's keep going. So basically when the uh robot be in contact with the object uh here we actually have three signals. One is uh the the the the robot state uh and the other one is the uh command for the uh controller. So if you equip the robot uh compliant controller uh most of the time you can see this penetration because uh the the arrow between the target and the state give us the force on the object and then if we have t sensor on the robot body we can estimate the result headed contact patch which is the red part here. So this is a very uh simple uh diagram in command control and most of time we can run PD controller on the drone. So the whole system is kind of like a uh this kind of like system. Um um yeah it's very common in uh control theory but here I think the good part we observe is if you actually equip the robot with T sensor and the compliance controller everything is measurable during data collection. So which means uh if everything is marable uh you can develop some like datadriven method to learn these variables and uh if you think about it uh so when we have this kind of like behavior during the compliance control this three signal actually define where is the contact on the hand because we have the connect page or the whole hand and how large is the force because uh I think this uh I mean at the at the one hand we have the sensing on the robot robot fingertip can tell us the the force. Uh on the other hand, because the arrow uh uh between the command and the the robot state, it can also give us some of the proxy of the title. Uh so I we believe like uh all these three variables are pretty efficient to define contact happening in uh data manipulation. So and as I mentioned since everything is meable during teleation uh our idea is convey just you know uh uh like develop a data collection pipeline which can collect everything like this three signals duration and develop a learning method that actually can ground all the information into this uh three variables instead of just katic targets. So that's kind of the the the the idea of the the the method. So to verify that we actually did a pretty simple experiment first. So um as you can see from this diagram uh in this experiment uh we have two input which is title and the robot state and uh as I mentioned if you equip the robot compile controller uh most of the time you need to have some penetration to like for the command target to make the contact actually happen. So we want to predict. So if we have these two variable, we want to design experiment predict where is the uh like the actual u u um target for the controller. So we basically did a like a grasping uh like a data data collection simulation. Basically the the human just try to grasp different object simulation like I show here like a apple or a box or some like random objects in simulation. And during during data collection we can have this kind of like pairs of for example the tail signals uh or the han and we can have the robbo state which is the the blue skeleton.
And we can also collect the uh controller target the the penetration state which is I think the uh here is a red script and uh yeah we basically collect some data this kind of data and train a very simple network doing this take two of them and predict the other and uh just to verify if we can do this on the unseen graspy uh and uh it turns out I think performs very well uh as you can see here and and we also did some very simple So like sweep on the architecture of the different backbone and different prediction paradigm like residual absolute uh and then we can see that this kind of like small network can transfer pretty well unseen grasping. So uh I I'm showing this figure this four grasp are totally unseen during data collection like unseen grasping kind of very random grasping um and uh if we just train a network trying to do this kind of prediction uh it can work pretty well. So if you say this this column um the uh the right part is the uh ground truth we collected from the data collection uh and the green part is our prediction. Uh so basically I think this this experiment uh back then give us a lot of confidence. Oh, if we uh if we so as I mentioned we think three these three signals can represent like any kind of like contact and uh this experiment basically verify that if we have two of them we can predict the the the other one because we think these three variables they are coupled and uh they are based on the the the physics when you have the grasping interaction and uh based on this is so this is actually super so this is really cool but I had one question about the predict so you're predicting the target state.
So, and you can do this with um with so so you said you could do this with with unseen previous unseen grasps. Do you think you could do this with like a like a glove like a tactile glove or something like that instead of with a robot gripper like could recover the so like you could recover it from from like more scalable data sources. Does that make sense?
>> That makes sense. So, so I think that's a very good point because uh um if you think about if we have some like a really really good tetlow u but I think the tagglow only can give us the the the act robate and the signal right. So >> exactly if you if you really want to train a very good policy from this kind of cloud data you need to think about how you can predict the penetration of the target right like the the control target. So, so in that case um yeah I think I think this definitely should I mean the so if you want to do this kind of prediction I believe the uh ground truth definitely from actual physical system because if you think about this kind of like prediction why we have this gap between the target and the actual state this is because the the low-level dynamics of the robot >> uh and uh we don't really have that but I think one good idea is as we show here you can train a really small network to make like unseen prediction happen. So which means maybe you can collect a small scale robot data trying to do this and then you have a large scale cloud data. You can kind of use this kind of small scale robot data to analy large scale cloud data and then later on you can leverage this kind of large scale data for poly training. So I think that's kind of that pretty much a very durable uh solution. Um yeah >> yeah okay that that makes a lot of sense. So and then the other question is sort of more of a clarifying one. So this is basically just um for anyone listening that like so the understanding here the reason that you need this target state is b is just because the tactile information plus the current state is not enough to fully recover the forces that you want to apply. Is that is that the right understanding right intuition here? Yeah. And so yeah okay cool.
>> Yeah.
>> Okay.
>> Right.
>> Yeah. So yeah I can move on. So basically uh so after we verify oh uh if we have this kind of representation of the contact uh you know it kind of like a transferable to unseen grasping and you can really train a very small network to make it happen. Uh then later on we were thinking about maybe we can just put it in the policy. So our idea is very very simple. You can think about like most of our uh network is just a very standard diffusion policy. There is no print train. There's no print trained encoders just everything t from scratch but I think the special part is we kind of add the previous uh like the three variable mapping stuff into our action hack. Uh so we have this like two faces training process. So during the first phase we just use diffusion to predict both the tactile and the the kinematic target which is the the actual robot state. uh I think nowadays it's very common uh uh for a lot of like I mean a lot of people did a similar stuff uh because if you have some of the modality in your uh trajectory in your uh data you pretty much can put all the modality as objective I think a lot of people also say this is a like word action model especially if you have you have the vision modality you can uh like forecast the future like vision d like v dynamics so here we have a similar idea so basically we want to we want to use diffusion to predate both the robot state and the titel field title uh and to achieve that it's not very easy we can just you know uh do some lat diffusion stuff to make the prediction very efficient um but uh yeah like I mentioned uh um and also Chris asked the question basically um if you only have these two variable and for if you input this variable to the uh like controller it cannot actually achieve the title you predicted so the reason people want to predict head is bas basically say uh you want to say what is the desired field title look like right u but uh if you think about it uh uh dreaming title is not enough we want to actually make make this headle be executable by the controller so that's why we have another mapping so this mapping is uh as same as the the the mapping I showed in last slide so basically we want to try to use the catal signal and the robbo state signal I mean although they are predicted from the diffusion we want to leverage those signal to map to a control target uh kind of like bridge bridge the gap from the prediction to the lo control uh and then uh input this control target to the compress controller. So uh in the end uh in our pipeline we will have three variable for the prediction. So one is the T signal like this and we have our like the the actual robbo state which is the blue skeleton here and we will have one more step try to predict the penetration part like the the actual uh uh robot target for the controller. Uh so that's a very simple idea. Uh and uh yeah like I mentioned there are two part the first part is we want to predict what contact the robot should create and then the second part is uh we want to bridge the gap from the prediction to the afro execution uh to learn about how to execute the contact.
>> I have a question. So >> Mhm.
>> Uh if you go back to the previous slide, right? I guess like my sense is that like the the tacttop plays a huge role because for example the wrist camera most of the time the fingers are oluded especially if you're doing nonprehensile manipulation. So is that something that you guys also felt through this uh I mean ablation or like you know through through the experiment you felt that actually maybe the the the perception part from the image is less important as compared to the tactile. the mo I mean the the perception tells you where to reach but then the tactile tells you how do you sort of like you know stay in contact for the task. Yeah, I think I think so. So, so I think uh so that's a very good question. I I think uh for sure I mean first point I do think this kind of task uh like a wrist camera is very important. If we don't have the wrist camera we cannot performance. Uh yeah that's the first point I that's first feeling I have this uh but the second one if you take a look at this image so this is from our simulation. So from the wrist you the the I mean you can tell most of the information uh for example if you want to reach to the object you can tell from the image but uh because of the occlusion everything you cannot really tell if the robot is being contact with the object uh and uh on the other hand you cannot like uh tell how much force you apply or not so I think because of this uh especially for this kind of like a text handulation uh tactile could play a more important role. Got it. Got it. Did that answer your question? Yeah. Yeah.
>> Okay. Uh yeah. So uh yeah this is a like more technical details. Uh I mentioned uh we use latent division lation for ta prediction. And I think one interesting observation we had is you can actually compress tactile images to very small or tail like in general t signals including both tail images and tile areas to a very small dimension and do very efficient prediction within this like a very small latence base. Uh so yeah I think uh this is this is not very surprising to me because I also had a paper previously showing a similar stuff. Uh so I think the reason here is comparing with vision. Um so tactile is definitely less informative because uh I mean it's is more informative for the detection style but uh it doesn't have all the background the color stuff it's just the the the the force distribution the contact geometry uh this kind of very very simple information so which means uh I mean even though we have this kind of vision based sensor uh which is kind of like uh provide us very high resolution images but we can basically compress this to a very small latence space and do the prediction there. So yeah, this is also I think one of the reason can make the our policy can run really well. Uh because if we try to predict into the original space, it could be more chaotic and more difficult. But if you can do this with diffusion stuff, it can make everything easier.
>> How fast can you do it? What what's the frame rate of your policy?
>> Uh I mean we cannot do very fast. It's kind of like so we can run on real int.
>> Yeah. So it it's not that efficient but I mean if you predicting the original telespace it will be like even less efficient. I think one of the reason is we also have cameras we are using two cameras and we use so uh even though we predict all the title images uh in the latent space we still need to encode all the title images as the part of the condition. So you can think so basically our system is running with six cameras four like fingertip cameras and two like RGB cameras. So in that case yeah because I think that that is also one of the bottleneck make our policy kind of slow but I think five herz is also pretty now for solve a lot of like reaching time.
>> Yeah I think five is pretty normal. So >> yeah yeah and uh yeah another part is I also want to introduce some of our like system engineering we develop a very cool you know data collection system can give can provide us very high quality uh you know dual location data. Uh so uh we use two different uh uh like a motion capture methods. Uh in the real world we use the like the mocap uh probably this is one of the most accurate system we can use uh nowadays and in simulation we use the um like meta quest. This is a pretty common basically use meta quest to track the finger and uh in our pipeline u our targeting is also very straightforward. Basically we want to map the the human fingertips uh with the robot fingertips. We do not do the human like a hand body retargeting. We we just try to map the fingertip from human and robot. And uh then uh after we map this kind of fingertips, we use uh an algorithm which called relax uh to solve the configuration of the hand. And then after we have the targets for the hand and the arms uh we designed two different uh uh compliance controller for the hand and arm to make the whole system compliant. For the hand we just use a joint space controller and for the arm we use OC. Uh yeah and uh I think as you can see in this video uh by doing that uh we can we can we can actually collect the pretty high quality uh data.
So we can uh actually control this uh like algorith doing this jaw opening task smoothly. Um yeah and of course for the whole process you need to tune a lot of parameters for the for the retargeting the tracking you need to make make sure everything works well.
Okay. And then uh I think here is an interesting ro video. This is from uh our policy ro and uh we did some like visualization of this uh like ro basically. Uh so we have as I mentioned we have our taup prediction and uh uh so the taup prediction is basically come from the previous time steps and uh after you predict the title you execute the targets you can observe the the f title in the in the coming time steps.
So we time align this prediction with the uh observe title uh like make them like align well. Uh so as you can see here um the differences between the observed title and the predict title are pretty small. So which means during so within our policy we are not only predict the title we actually kind of make this tail executable. So as you can see here during the wiping process when the sponge being contacted with the dish we will have some some of the you know large force on the fingertip because the robot is trying to apply force on the dish and uh we can also tell from the skeleton uh I show it here. So when I think yeah I so basically when the robot be in contact with the uh dish you can see that the difference between the this like a uh green hand and the green skeleton and the blue skeleton not not only for the finger finger side also for the base frame set. So this is because the robot arm is also trying to apply a force on the dish. Um so yeah I think I think uh this video basically shows our algorithm can run as I as our as as we expected. We have our prediction and we we can make our prediction executable during lot and uh yeah this is the second task the the ad grasping we also have the signal on the fingertip and the prediction and observe talign very well and we have the this third one which is the box flipping. So I think the boss flipping sometimes you have some mismatch between the actual prediction and the observed tactile. For example this frame they looks pretty different.
Um but I think yeah and uh most of the time frame the the title pattern looks pretty align like this. Okay. And uh yeah then we uh evaluate our policy on five different tasks uh including um like nonpre manipulation like cross flipping and like opening tasks as well as the fragile object grasping like the fragile eye grasping. So we compare with uh because like I mentioned our backbone is very simple just different policy. So we direct we just compare with two different policy uh uh like type base uh they are both different policy but with different like conditioning. So one is just a pretty standard visual motor diffusion policy and the other one we add tactile as additional uh additional conditioning for the diffusion policy and uh among all three all five tasks uh our method achieved the best performance.
Uh yeah and uh I also want to show some of the behavior learned by the base and our method. So this is a wiping task. So as I mentioned for this wiping task during different ro we are actually changing the the height of the table to make this task like challenging. Uh and you can see here for our t for our method it can do this wiping smoothly.
uh but one one of the typical failure mode of the baseline is the policy tend to kind of like overfeit into the kinematics and then get stock around this dish. This is why our policy achieve better perform than baselines and we also demonstrate the some some of the comparison between our policy and the uh baseline on this flipping task.
Uh yeah so so as I showed before for the baselines I think one of the the common failing mode is the policy cannot provide enough force for flip this box uh because I think this process definitely need the the penetration I mentioned you need the gap between the target state and actual state >> I have a question so what's the difference between the left contact grounded policy and the right contact grounded policy >> you mean this this two >> oh they just uh like same policy but just we just test the different position of the box different.
>> Oh, okay. Okay. Okay.
>> Yeah. And also the I think this is the the comparison for the job opening task.
I think one of the failure mode for the baseline is sometimes it just cannot provide enough force or you know um like like after a while the the the jelly lead kind of get stuck and you cannot really open it. But for our policy yeah if you give give the time long enough it will eventually open it. But uh yeah this one is already yeah this one still get stuck here. So okay uh so yeah basically this is our takeaways. Um uh so uh to summarize contact grounded policy has three key takeaways. Uh so first uh oh sorry uh so it represent so in our method it represent contact for multif finger hands using variable three variable signals the actual rubal state sensing and the control target. So this is very common in control theory but we found is very useful and we kind of like uh bring it to the like supervisor policy learning. Uh second uh we uh like predator method to bridge the kind prediction and and execution by converting the predicted tail and the robot state into a executable uh uh compliance control target. So uh in this case the predicted tactile the predicted contact realized by the lo controller um in this case uh uh so we we not only dream the f tile we also make the talle dreaming uh like executable uh finally uh all this design leads to a better performance comparing with baseline uh I think we have around like 15% uh policy improve performance improvement uh yeah uh and I I think that's pretty much the the printation.
>> Really cool. Thank you. Um but one thing I was thinking of while you're looking at this, so since you are predicting these compliance targets, do you did you think about it all about like what are the gains? What's the controller that's trying to reach those targets? Does that something that you think you'll want to predict in the future or >> is that not necessary? Yeah.
>> So that is a very very good question. So I think one limitation of current method is uh because uh we kind of trying to represent contact with three signals right the target the robust state and tail but if you think about it this this is actually conditional on fixed games.
So if you change the game maybe I mean if you have the similar uh like a controller target or the or the similar uh robot state maybe the t sensing will be different because the force will be different right so that's why all our current method or experiment only evaluate on fixed games so we won't like as long as we collect the data we won't change the game of the controller to the robot during deployment uh but in the future I think if we really want to make this kind of method scalable we should design some method conditional control So >> or even condition on the type of controllers. Um if you think about uh I mean I mean on the one hand I do believe for different body different hands they do have some of the uh like similarities for example they are I mean they all look look like hand right which means their dynamic system should be very similar. If you think about uh we can collect data across different like hands and we maybe also data about their gains. there are some like dynamic system parameters. Uh maybe we can train a more generalizable mapping uh not only condition on cat sensing and robust state but also condition on all the dynamic parameters and the control gains as I mentioned.
>> Yeah. And I um I think that uh um a very similar question uh is that uh control gains and also whether uh the object it's soft or not or or the softness of the object that also affect the mapping and that's also not something we are not considering in this paper because in the paper all the object they we assume they are rigid and uh they they are not soft but if they are deformable then I think uh the the policy definitely need to condition on more things.
>> But wouldn't it so you have the image of the object, wouldn't it learn? Wouldn't you expect it to learn like how soft the thing is from demonstrations or >> Yeah. Yeah. But uh but I I think pro probably v vision is not enough. For example, there is a ball. I I don't know how much air in the ball. So I don't know how. Yeah. So So yeah, >> see history basically like v video visual tactile history or >> Yeah. Yeah. Probably the policy needs to try it. I I I I don't I I I I haven't think it through but I I was just going to say that uh like GANs and the uh softenies of the object they are they kind of coupled stuff and both of them can affect the mapping the contact grounding mapping.
>> Yeah.
>> Yeah. So I guess what uh like if you had another few months to work on this what would you do? What's the next steps for this project?
>> Yeah. So, so the I think the the the thing is at the beginning uh about at the beginning I guess the the the project is like how can we leverage BL data for learn policy uh and uh yeah and I think the the thing is nowadays I don't think all the glove uh so back back then when we intern like too many tea glass for us to use u that's the reason we shift to this project but I do think the the point you mentioned is also the idea we want to work on in the future for example um it because and I so there's a lot of like cat nowadays and I believe most of them uh they try to learn the policy from the kinematics reading from the glo uh but the thing is like if you really think about the interaction for the uh controller if you only have kinematics it's not so when we uh for example when we grab this phone uh I mean if you have the blow you will record or you signals on the fingertip and you have this kind of like hand skeleton for bus 8. But uh when we when our human grate our intention is not like trying to make our hand close to this kind of like configuration but we actually trying to make our hand penetrate inside this phone that's the control target right but I think most of the uh work working on like a learning policy from glow they kind of like miss this part uh and uh I do believe if we use a similar idea uh for this kind of like learning policy from large scale glo data uh for example we how our kinematic target like kinematic configuration like the hand configuration here but we also try to u calibrate or try to map the actual uh like target the actual intention. So which should be a like like smaller hand like a penetrate in within this form. So in this case I I believe we can learn better policy because uh you actually consider the gap from the kinematics and the low-level control execution. Yeah.
And also I think uh this kind of idea also works in uh yeah I'm actually recently I'm working more on the side and I found that if you have this kind of like um compliance uh controller and low level and you have this like target stuff it will also helpful for RA especially for the inhand rotation in handling stuff.
>> Yeah that makes a lot of sense. I wonder because like like you're saying like the tactile feedback doesn't exist for the teleop glove. So like this may usually so that makes sense that RL would maybe work really well with this. Is that kind of your Yeah. Yeah. Cool. Uh because just because like obviously the demonstrator is not going to take advantage of it. Cool. Okay. That makes sense. Um you're going any thoughts? You muted. Sorry. You're muted still.
>> Sorry. Uh uh uh Jon uh can you go to the uh p pipeline slides? the pipeline of the work.
>> Yeah.
>> Yep.
>> Yeah. So, uh basically the current pipeline involves two two big two policies, right? the diffusion policy and also the the mapping the contact grounded uh mapping and uh so um um so when we start the project we were uh thinking something ambitious like like for uh for example for the same policy sorry uh for the same task we maybe we can keep the diffusion and uh uh change the low-level or the contact consistent policy um sorry the contact consistency mapping uh if the controller is changed or the the PD g of the controller is also changed and uh um vice versa. So uh we may able to train uh fundamental contact uh contact consistency mapping and uh for and that can be used for multiple task. So for different task we only need to change the diffusion. So basically um in that way uh the the contact consistency mapping is more like a foundation model for contact or something and that can can be generalizable across different tasks. Yeah. But uh we didn't have enough time to collect enough data to train it. But I always feel like um if we can show that in experiment that the two policies they can actually uh uh like like um plug into to different module and that that can be really cool.
Um I don't know whether I I make it clear or not.
>> Yeah, interesting stuff I think.
>> Okay. Well, I think uh we can I guess that this is this sounds like really exciting stuff. Um thank you guys for being here. I guess we have one last question that we usually asked to wrap wrap up which is um basically any other cool work that you'd like to give a shout out to. Anything you saw recently that you thought was particularly cool or Yeah.
>> So I think the because I'm working on the IO style this so I think this year uh I said there's one paper called sim to real that is pretty cool. Uh yeah and uh yeah I think I like the so I think the one one thing for current IO is you can train very good IO policy with present information for example the object post stuff um but in reality if you have those kind of like policy condition object post is very to very hard to deploy as well as very hard to make it useful. Uh and I think to reality they find a very good way to make it useful. For example, you can, you know, actually use some of the foundation model to get the post object and then, you know, try to uh like get a sequence of the pose from the video and then um because their post is pretty generalizable. You can just condition the video and the post from the video and make the the actually uh you know uh like work for some real tasks which is pretty cool. And I do think their their IO is uh is really um like uh impressive because this is I I was working uh always working on like reward tuning and how how you can you know design good rewards to make the IO scalable. You can paralyze the a lot of environments and also can get some some of good policy. I think they really did very very well on this part. So really impressive >> and um I was mainly uh reading papers about system zero of human noise like like sonic the the famous paper and uh uh this might be relevant to something Jungong just said and there is another paper called B B B B B B B B B B B B B B B B B B B B BMF0 which use unsupervised RL that can avoid the the painful process of reward tuning and stuff. And uh I found that idea it's uh I mean uh uh it's really uh appealing but I haven't tried it myself so I don't know how how good it it work in real life but I really wanted to try it. BMF zero.
>> Okay. Yeah. Yeah. BBM BFM0.
>> Good. Okay. Cool.
>> Yeah. Well, all right. Cool. Well, thank you.
>> Uh well, thank you guys for joining us.
I think it's really cool work. I really like the lots of really nice results of rearranging objects and thank you for coming on hopefully the time.
>> Yeah, thank you for having us. I'm very excited to share our work.
>> Yeah, thank you.
>> Thank you. Thank you guys.
Related Videos

Setting up a curved screen with Immersive Calibration Pro 4 and multiple cameras (P3D v4)
FlyerOneZero
23K views•2019-07-21

Robot Learning with Sparsity and Scarcity
allenai
379 views•2025-10-14

Jorge Mendez-Mendez: Unlocking Lifelong Robot Learning With Modularity (2023-10-05)
umassmlfl
237 views•2024-01-06

Northwestern’s MS in Robotics: Student Robotics Projects, 2023
NorthwesternEngineering
1K views•2024-05-31

"Perfect" Turns: Turning by the Gyro - FIRST LEGO League (FLL) SPIKE Prime + EV3 RePlay Programming
ZacharyTrautwein
94K views•2020-10-02

Gorkem Secer: TSLIP-based Deadbeat Running Control of Bipedal Robot ATRIAS
DynamicWalking-wv6qm
298 views•2018-06-22

Self-Driving Cars Need Lessons On Human Drivers | Maddie About Science
skunkbear
26K views•2018-08-21

Milrem Robotics’ THeMIS UGVs used in a live-fire manned-unmanned teaming exercise
MilremRobotics
99K views•2021-05-20
Trending

we're almost finished the house (ep.125)
JennaPhipps
347K views•2026-07-22

We Finally Know Where Saturn’s Rings Came From
astrumspace
79K views•2026-07-22

BIG BET: Cathie Wood goes ALL IN on Elon Musk
FoxBusiness
89K views•2026-07-22

MIC DROP: Smithsonian Director Called Out For Woke Propaganda
TheAmalaEkpunobi
37K views•2026-07-23