This talk presents approaches to overcome two fundamental data challenges in robotics: (1) data sparsity, where tactile sensors provide only local sparse information about contact points, addressed through model-free reinforcement learning with active exploration policies that judiciously select next actions to maximize information gain; and (2) data scarcity, where collecting biosignals from disabled-bodied subjects is extremely challenging, addressed through generative AI methods like CHIMG that leverage vast repositories of previous data to synthesize additional training samples, enabling intent inference with minimal real-world data collection.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Robot Learning with Sparsity and Scarcity
Added:to Jinxy to um share his content.
>> All right. Yeah. Hi everybody. Um and we are coming here for presentation by Jinxi. Uh he did his PhD at Colombia work co-advised right I think by both by Shiran Song and Mate Chocoli who are both like leaders in the area of robot manipulation learning uh doing amazing things with with manipulators in general. Um, Jinxier has done some really exciting work in that space also looking at um, tactile information and how to do control and manipulation with that information. And he is currently at you at Boston Dynamics.
>> Uh, the ray dynamics >> at this >> robotics the robotics AI institute. Is that the right saying?
>> Yes.
>> Because then it should be >> robotics and AI institute.
>> Oh, so it's Ray with a double I. It should actually So these guys are I think as confused as I am about their naming, but um uh and yeah, so I'm very excited to see you see your presentation.
>> Thank you. Thank you for the kind words and introduction. Uh so I'm going to go ahead and share my screen.
Um, all right. Can everyone see my screen?
>> Yes.
>> Can you guys see the slide? Uh, the first page of the slide?
>> Yes.
>> Okay. Um, so yeah. Hi everyone. I'm I'm Jinshi and um um I'm very excited to to be here to to present uh for AI2 and thank you all for taking our time to to drop by in on on a Friday afternoon before the long weekend. Uh I really appreciate your time. Um so for maybe for small clarification questions we can uh ask as I present but for longer discussion questions we can leave it at the end and also um because I I can't see people's face I can only see the pre presenting screen so if there's please stop me um if there are any problems or questions okay so I'll get started um so today's talk I want to talk about some of my previous work in robot learning uh specifically how to address the sparity and scarcity issue in uh robotic data.
So I want to start with um um some of my own experience how how what my own journey is entering robotics. So when I first came to Colombia I didn't know that I want to work in robotics. So I worked with different groups on pure learning and pure vision problems. Uh and then I started collaborating extensively with professor Peter Allen on a wide range of different applications. So I work on motion planning uh n visual navigation. I also work from uh dynamic grasping to um I versus control.
And then I realized uh I think one of the key challenges um which make you robotics so unique so different from vision and language uh is the challenge of data which is the lack of data uh we have in robotic applications which I think most of you guys will agree with me and I often time use this spectrum to uh to to see how much data do we have for different AI applications.
Um and on the left side is the data sufficient uh is is data sufficiency and on the right side is data insufficiency.
Now for language and vision applications um it's definitely uh on the more sufficient side with a relatively more sufficient data we can use especially from the internet and that's why we have amazing applications such as chargebt uh uh Darly um and we move to robotics things become uh trickier so in robotics often times we need specialized or customized hardware to do the job. And this hardware requires a specialized knowledge or expert domain knowledge to operate. Meaning that it's really hard to collect data at large scale with this device. Uh but luckily sometimes we we have simulators and I'm a huge believer in simulation. Um and if we formulate our problem smart enough um and then it becomes simulatable then we can actually enable large scale training inside simulators.
However, for some applications in robotics simulation is very hard. Uh for example to simulate raw tactile signals or to simulate granular media uh it's very tricky to to to have good simulation for these applications. So it's more trickier. it's slightly trickier to to to do data collection.
Uh what's even more challenging is when you want to collect data from human subjects and in these applications such as human brain signals or application that use muscle signals, there's no simulation at all. Um meaning that it's even harder to collect data in these domains. And further to the right, um the the the most difficult case for data collection I believe is when you want to collect data from subjects with impairments or disabilities.
Um for for you know it's already hard to to collect data from human subjects let alone uh collecting data from people with disabilities.
So I think that's the application domain where data collection is the most uh challenging.
And often time when we look at the data problem in robotics we we look at it from uh the quantity of the data from that angle and that gives us a data scarcity issue meaning the sheer amount of data we can collect is very limited.
However, there's another angle to look at this problem which is from the angle of the representation of the data or the quantity of uh the quality of the data and that will give us the data sparity issue and that means there's a large proportion of empty space or nonrelevant information in the data we collected and this is because robotics uh is the science of interaction. So we have a so a piece of software or algorithm embedded in a piece of hardware and the hardware is making active interaction with the world to collect useful information useful information and then the the gap or the less between two con consecutive actions makes the robotic data inherently very sparse. Um so in this talk I want to focus on two domains. Um so I want to talk about selective works from tactile manipulation and rehab robots because I believe these two domains are perfect examples of how uh robotic applications suffer from um data sparity and data scarcity issues. So tactile sensing is a very powerful modality. Um but at the same time it's challenging to use because it's active sensing modality and every time you contact the target object you only get sensing around the local area of the contact which makes tactile signals such as locations and normals inherently very sparse.
Now for rehab robot as I as I said before um it's basically I think the difficulty of data collection to the extreme. So I've been working for a long time with uh stroke patients and it's really hard to collect uh to schedule session uh with clinical experts with stroke patients to to collect data and even for the precious data collection time there's a huge setup time and patients get fatigued very easily. So it's very very challenging to collect enough data to to to make rehab robot work.
So this will be the two domains I want to focus on today and I will start with uh the first domain tactile manipulation and I'll talk about how we address the data sparity problem by utilizing simulation and model free RL to learn efficient tactile only um exploration and manipulation policies. So in this line of work we go from 2D problems to 3D problems. we go from identification to manipulation and we eventually go from above the surface uh to under the surface.
So I will start with due to time constraint I will start with tendon 3D which builds on top of the code training framework uh in tendon and extend this to 3D uh space.
Um so this is inspired by humans capability of using tactile sensing because we we humans are experts at using tactile sensing. We can uh we can put our hands into our pockets and retrieve a particular object without looking into the pocket such as a coin. Now we want to ask if a robot equipped with the tactile sensor can it do the same. So we are trying to identify a particular object um with as few actions as possible. So not only we want to identify the target object but also we want to do this very efficiently.
So that's the goal of the of this project.
So before I go into the setup of this particular project, I want to take a quick detour and talk about the finger, the tactile finger we used uh in the tacttop projects that I'm going to talk about today. So this is the disco finger and it's developed uh in our lab. uh it has uh light emitters and receivers on the skeleton and then when you touch the gel you're going to change how the light transmit inside the gel and consequently change the light received at the receivers.
So there are 30 pairs of light emitters and receivers. So this gives us 900 dimen around 900 dimensional raw tactile signals. And this is a visual of how the signals look like when you touch somewhere on the finger. And then we can train uh a machine learning model that can regress from these 900dimensional signals uh into uh contact locations and contact normals and contact force. And this regress information will be the information to the policies that we we'll talk about today.
Okay. So going back from this detour to uh this particular op uh project. Uh so we mount this tactile finger on the UR5.
Um and a key advantage of this tactile finger is that it has allaround sensing coverage. So uh it can sense touch from uh all directions and then we'll put this target object which is unknown and in a random position and rotation uh in the workspace and then the robot basically drives the tactile finger make poking the target objects collecting a sequence of contact locations and normals this sparse information and then try to identify this object out of 10 YCB objects.
Um, so here are just four examples of the information that the robot collects.
So I have a very quick uh question to the audience. Can you identify these objects based on this sparse information? Anyone wants to give it a try?
>> So the red dot is the touch point and then the green is the normal.
>> Correct.
>> Right. Yes. Yes.
I can't >> is the Which one?
>> Third is the >> the third one is the drill.
>> Drill, I think.
>> One is a drill.
>> It is a drill.
>> Good job.
>> Okay. Anyone else? Or I'm gonna review the answer.
Uh the first one is the stand.
Okay, I'm going to review the answer.
>> Wait. Okay, so actually you guys are actually right. The third one is the drill. Um I wasn't expecting you guys to be correct, but uh but okay, the point here is uh I want to show that with very little information uh it's very hard even for human to identify these uh objects. uh but tenant 3D can recognize these objects. You can only touch with very few actions. So how did we achieve this? Again the challenges are uh tactile sensors are uh active sensing modalities and they only gives you very local and sparse information.
So in order to approach this challenge we uh we uh design we formulated the problem as a model free error problem so that we can learn an efficient exploration policy to judiciously choose the next action.
So in our formulation we design a six action space. So at each time step the robot can choose from XYZ ro to either increase or decrease for a small step.
So this gives us 12 discrete actions and then we identify three key components for this task. Uh the first one is encoder which encodes the sparse information into a 3D representation and then the discriminator will look at this representation and predict the identity of the object.
And then finally the explorer uh will look at this representation and uh guide the exploration will output the next action to move. So specifically um the fingers contact the target object. It collects a sequence of contact locations and normals. We first rearrange them into a point point cloud and then this point cloud is passed through 2.8++ architectures to generate the 3D representation.
We choose three uh point plus because it's permutation invar invariant and it can also handle different length of points and this representation will be passed to the discriminator. It will predict identity with a confidence estimate and if the confidence is larger than the threshold uh it will terminate this episode with the final prediction.
Otherwise, the explorer will look at its representation and compute the next action uh for the robot to move.
Despite that, we have a a distinct uh discrimination uh decision- making policy and uh exploration policy, their training is actually highly interlipped.
So we modeled the discriminator as um a CNN and trained with supervised learning and we model the explorer as I agent and trained with PO. So during this co-raining process the discriminator will provide termination signals to the explorer.
uh it will terminate when the confidence is large enough and then when we are training the explorer it will collect a label data set of the point clouds with their ground truth identity and this label data set will be used to train the discriminator.
So during this co-raining process they co-evolve they they uh change each other and then they eventually converge and the whole pipeline is trained purely in simulation and then seem to transfer onto the real robot and let's show some of that uh policy in action. So the object is placed in an unknown position.
Um and then the finger will visualize the predict distribution over the 10 objects. We'll also visualize the point cloud um representation and the finger um collects a sequence of sparse information and it gradually moves to the most discriminative area of this object. And for this object is the the picture handle. And through a small swing motion uh it gets a very discriminative point on the handle of the the pitcher utilizing the full sensing coverage of the sensor. And after another discriminative point on the handle, it quickly converges to the ground truth um identity of this object.
And this time it takes 37 actions.
And this is another example on the power drill. Uh so this time it gets a very discriminated point from beneath the drill and then it quickly converge to the ground truth identity and this time it takes nine actions.
Um so to summarize this part uh yes we have the data sparity issue in tactile sensing but tenant 3D utilize active exploration to tackle the problems by meticulously choose the next action to take to collect the most useful uh tactile information.
So that's the tenant 3D work. Um now we want to uh move to a uh even a challenging work a more challenging environment uh where vision is basically out of the question. So this this work is we what we call geotech um and in this work we are trying to retrieve objects buried completely inside granular media with only tactile sensing.
Again we still have the the challenges of local and sparse information of tactile signals but also uh because of granular media in such a sensor deprived environment we also have extra large uncertainty and sensing noise. In addition this environment also requires certain properties of the hardware. A lot of the uh vision based camera based tactile sensors uh the pioneer work even though they have very high resolution they have boxy form factor um which will induce high drag when they move inside granular media also um um the the tactile sensing they gave they don't cover the the edge and corners. So but in granular media you really need the full sensing coverage. So there are some specific hardware requirements on the sensor we use.
We borrow the idea of pushing from uh a lot of previous work. So many previous work show that pushing is a very efficient technique to reduce uncertainty uh and uh increase grasping robustness.
But most of them show this technique on tabletop or open air environments. Uh in this work we basically argue that uh pushing is efficient techniques also in granular media. It reduces uncertainty increase robustness. It also is a great technique to follow the object to a grasp of pose.
However, designing a um heristic based pushing behavior isn't uh trivial. So again we formulate the problem as model free RL and utilize simulation for large scale training and we um and we hope that the pushing behavior can emerge by itself through the design of the action space. So our action space is a fourdimensional vector um while the first three elements is the discrete change in x y and theta direction. the our finger because our finger moves in a 2D uh plane. Uh and then the last elements indicates um whether the the exploration should be terminated and the finger should close and grasp the target object.
Um any questions so far? Because I I can't see people's face. So >> no.
>> Okay, great.
>> Sounds good. Um so >> perfect. Uh so I'll continue. Um so the observation to our policy is a sequence of contact location contact force and past actions. Um however different from open air environments in granular media environments there are uh actually two different kinds of contacts. The first contact is what we call ubiquitous contact and this is the contact you get from um between the finger and the granular bees.
and before your finger started pushing the target objects. And the second type of contact is what we call pushing contact.
And that's the contact you sense when the finger has started moving the target object. So we believe only pushing contact carries important information for our policy. So we use a futuring techniques to only keep pushing contact to our policy.
And also very importantly we have a curriculum training strategy.
Again the full uh pipeline is trained in simulation and sim to real transfer onto real robot. So we want to reduce the sim to real gap. So we think it will be very useful to train in simulated granular environment to simulate the behavior of the target object when it moves inside granular media. But again like as I we briefly mentioned before simulating granular media is very time consuming uh very inefficient. So we want to reduce the time in granular media by bootstrapping our training on a tabletop environment. So in this curriculum strategy we first train a push to grasp behavior on a tabletop environment and then after it converge on tabletop. So basically the finger just push and form a group grasp and and lift the target object and after it it manages this task we load the converge the checkpoints into simulated granular environments uh to fine-tune inside granular media environments and then eventually after it converge in granular media environment will zero transfer onto the real robot.
Uh so this is our curriculum strategy.
Finally this finally this is our hardware. We mount the disco fingers on the parallel gripper. Again, the full sensing coverage of the disco finger works really well uh with the granular media environment and the cylindrical shape also drags reduce drag inside the granular media.
I want to show some of the the policies in action.
Well, in this demo, we'll visualize the contact location and also the contact force.
So this is where we bury the chocolate box.
So our policy initially detects a fake pushing contact even though it hasn't really contacted the target object. So it starts to react to this fake contact until a true contact happen and then it starts to correct its own behavior um and then through a series of interaction uh it correctly uh form a good grasp and retrieve the target object.
So this is another example on the screwdriver.
So the screwdriver is actually a very hard object to grasp. It's very long and a very thin. However, the policy through again through a series of uh pushing interaction uh form a good grasp and successfully retrieve um the screwdriver.
When we train the policy in simulation, it's only trained on seven objects with primitive shapes. However, the policy actually works on a larger range of objects. So, these are some random objects we we chose from our lab. Um, and we just threw them into the granular media and then we asked the policy to retrieve them from the granular media.
And the policy is is able to generalize to these objects within three attempts.
So sometimes it does use up all three attempts to be successfully grasping these objects.
It's always fun to look at failure cases. The one of the most common failure cases is when there's a thin layer of granular media getting stuck between the target object and the finger. Uh and because of the very low friction of the bees and and and 3D printed object, it often slides off.
Another very common failure case is when we are trying to grasp cylindrical or ball-shaped objects because our finger is also cylindrical. So imagine two-cylinder grasp ball um is very hard and the ball is often times squeezed um uh outside the finger. So we believe this is uh a limitation of the parallel grippers and if we have multi-handed fingers hand grippers this will be less of a problem.
So to summarize this part, uh, Geot utilize model free RL and simulation to learn an emerging emerging pushing behavior to efficiently explore and grasp parag object um on sparse tactile information.
Nice. So we covered um my line of work um in tactile manipulation. Now I want to switch gear. uh feel free to ask questions if you have. Um now I want to switch gear to rehab robots and I want to talk about how we handle the data scarcity issue in rehab robots. So while data sparity in tactile manipulation focuses on the uh quantity of the data now we want to switch to the quality of data.
So in during my PhD I collaborate with uh the Columbia Medical School and um mechanical engineers. So our lab developed a wearable device to help the stroke subject open their hand. Um and we have to train machine learning models that predict their intent such that the the wearable device can provide appropriate assistance.
In this line of work, we explore different paradigms that ranges from semisurized learning to male learning to generative AI. We go from only able only being able to handle interession drift to also being able to handle interession drift. And I'll explain the difference in these two terms uh later in the talk.
We also go from being only being able to use low capacity classifiers to finally being able to use hyper capacity classifiers.
And the work I want to focus on in this line uh in this domain is CHIMG but I'll also touch on the the two projects before CHMG in this talk.
Again, the goal is to train in class with very minimum data collection because it's really hard to collect data from people with disabilities.
Before I go into details, I want to share some background about this domain.
Um, so 800,000 people have stroke a year in the US and a majority of them will actually be impairments. Less than half of them will recover full function. A lot of tasks that we healthy subjects we take for granted are impossible for stroke subjects and this is a stroke subject atte uh attempting to take and place this wooden block which is very hard is impossible. Um so our lab developed this wearable device on the right. Uh it's a it's a motordriven um exoskeleton. It has a motor. uh is a tendon-driven device and it also has a EMG armband. So the EMG is short for electromyioraphy.
You can think of it as you're measuring your muscle intensity or muscle activities.
The goal or the machine learning job, the software job is to develop machine learning models that can predict the user's intent from the sensor reading.
And if the model thinks they want to open their hand, the motor drives the tendon and open extend their fingers.
And if the um model thinks they want to close their hand, the model release the tendon and the patient can use their own strengths to close their hands. So the population we study uh in this line of work, they lose their ability to open their hand, but they still have their residual motion to close their hands, which is a very common uh stroke after effect.
This is uh the armband we use. Uh it had it has eight electrodes surrounding the forearm producing a channel EMG signals at 100 hertz. This is a demonstration on healthy subjects. For uh stro patients the change in the EMG signals is much subtle. That's also why it's much more challenging for stroke subjects.
We already know that it's really challenging to collect data on people with disabilities. And what makes it even more challenging is there are huge variations or drift in EMG patterns across conditions, sessions and subjects.
Uh so what what do I mean by these different terms? Um so a condition is a basically a combination of um the arm position and the status of the device.
So for example here are three different conditions um arm on the table or device off or arm off the table device is off.
Um and a session is a single use of device between doy and dolphin. So basically a patient joins our session uh come to the hospital we put on the device we do our experiments and they take off the device and that's a single session.
It it's easy to understand because for different subjects uh due to different neuromuscular impairments they demonstrate very different EMG patterns but even for the same subject if they come in different days for different sessions there's a huge drift in EMG signals as well waiting sessions uh across sessions now even for the same subjects for the same session if they move their arm to a different location there's also to the huge EMG signal shift.
And here's an example of for different conditions even for the same intent. You can see how different the the EMG patterns are even for the same intent.
And this is uh more pronounced for um the drift across sessions and subjects.
Um for the drift that happens within a single session, we call it intraction drift. And for drift happens across sessions we call it uh inter session drift.
So >> normally yes >> sorry just a quick clarification question. The different lines in each of these plots are these different runs or different dimensions of the EMG signal.
>> Yes these are different because we have a a channel EMG sensor.
>> So yeah so the x-axis is time. So the yaxis uh is basically in intensity and different lines are just different electrodes.
And if you would do multiple repeats of let's say arm on table device off there would also still a huge variability among those right or would they be >> yes >> yes >> okay >> which is naturally like you will have different patterns for different intents >> so here I'm showing like even for the same intent if the arm is perform performing the intent moves to a different position like the signals the pattern will also Does that make sense?
>> Yeah.
>> Okay. Uh thank you for the question.
>> Right. So uh because of the huge variation, normally the model does not generate generalize. So when we train a model on a single session or patient, it does not automatically generalize to a new subject or patient. So what this means is every time a new patient joins our session we have to implement the full data collection portal again which is a lot of burden on the show subject. So in order to relieve this uh problem the first paradigm we turn to is semi-supervised learning and the idea is really straightforward which is we want to utilize unlabelled data. Now in this work we basically only collect a small data set from the new session or patient and then we train a initial model we deploy the model and then we ask the model to adapt to new unlabelled data and this means that we'll need a huristic to label the unlabelled data and we don't want I don't want to go into details here basically the the heristic we use is a discriminant based semis vision and we by utilizing unlabelled data it does work better compared to baselines that do not utilize unlabelled data.
However, it has the its own problems.
So, because we are relying on huristic, this huristic does not generalize to the drift across sessions of patients and also this supervised learning um pipeline requires the classer to be able to run fast online update. That means uh it does not work with any type of classifiers especially for high capacity uh models.
And then the second paradigm we study is metarning. So in this paradigm we formulate the whole interferral problem as a multitask learning problem and we treat each session and subject as a single task. The goal is to learn some intermediate representation that is fatally adaptable to a new task.
uh in this way it does generalize to new sessions of patients but again metalarning has its uh limitation which is it requires the tracking of higher order gradients and it also only works with gradient based models. This means that it it this method is not classifier agnostic.
So at at the end of the day we we really want a paradigm that is able to generalize across different sessions of patients and also can work with the different type of class any type of class.
Um that's why we eventually turn to genative AI methods and that's why we uh study CHMG.
So what is CHMG? Chg is an auto reggressive generative model that can generate synthetic EMG signals condition on prompts. Um similar in concept to language model where a prompt is an existing piece of text in CHMG a prompt is an existing piece of uh EMG signals.
Um and then language model gen general gen generate um the next um uh words or tokens autogressively. CHMG will generate the next EMG a channel EMG signal at the next time step autogressively.
The way to use it has three stages. In the first stage, we train TMG on a large corpus of offline data collected across different subjects, sessions and conditions. And we hope that through this generative training, it can learn the general patterns of in different intents such as open, relax or close.
And now in the second stage, when a new subject joins our session, we don't run the full data collection protocol again.
We only do uh very limited uh data collection from this new context and then we keep sampling prompts from this limited new data and conditioning on the prompts. CHMG generate synthetic data by expanding the prompt to a length where the classifier takes and then this synthetic data is combined with the original data to train basically any type of classifiers and then we can deploy the classifiers on the device to to help the stroke subjects.
And in this way it can generalize to new sessions of patients. And because the synthetic generation process is completely separate from classifier training, we can work it can work with any type of classifiers.
So finally I want we we take the whole pipeline of data generation uh limited data collection into a real clinical environment.
Um and in this demonstration we are showing on boarding a new subject. So this the data from this subject is never seen by CHMG before and then we asked her to perform P and place uh because without any assistance it's impossible which makes a lot of sense.
Now we started data collection. So the way we do data collection is we'll give out verbal cues to the subject. We'll say open, relax or close and then we'll record the eight channel EMG signals simultaneously.
So you can see now we say open and because she couldn't open her hands so you barely see any change in the EMG signals for the open tent.
Now we say please close your hands and you see that she can close her hands and there are huge EMG changes.
And finally we say please relax your hands and that's the only data we collect. There's one round of open, relaxed, and closed data.
And then the first demonstration we show is on the cluster trained only with this limited data set. Um, and then because the data is so limited, the cluster never correctly predicts the open intent and the patient couldn't use the device to perform uh pick and place, which is very understandable because the data is just so limited.
Now this is the data we collected before with CHMG. What we can do is we can sample prompts from this very limited data set and we ask CHMG chg to complete the prompts.
And now the second cluster is trained with this uh synthetic data plus the original data. And now you can see the device can successfully open and the patient can use the device to perform multiple rounds of um pick and place motion and it's not perfect but it does work much better compared to the the previous model that do not use synthetic data at all. And uh to the best of our knowledge, this is the first time a model trained partially on synthetic data is deployed on a real stroke subject in a real hospital environment.
And to summarize this part, um in order to handle the extreme data scarcity issue, uh we use synthe data generation.
So, CHMG uh the G of it is that it leverages a vast repository of previous data through generative training but it also remains context specific via prompting.
So to summarize uh what we talked about today um we talked about how we handle the data sparity and data scarcity uh in two different domains. Um so we learned that even though we have the the data sparity issue in the tactile signals uh we can utilize simulation and we can utilize um uh we can utilize RL to learn efficient exploration policy to judiciously choose the next action to explore and to manipulate the target object. And for rehab robots, we learned that we can use uh we not only can we just uh try to make it work with minimum data by utilizing uh semisuralized learning or m learning, we can also dream up our own synthetic data through generative AI. Um the hope is that um the lessons we learn from these two domains are also applicable to other domains because the data problems are so fundamental in robotics.
Um at the end I also want to quickly um mention another um line of my work which I spend a lot of time trying to make it work is dynamic grasping. So we try to um make the robot grasp dynamic moving targets through online planning of grasp which are both motion aware and uh reachability aware.
And in some follow-up works, we uh study how to learn a meta controller that dynamically tune the meta parameters of each subcomponents in the whole data uh dynamic grasping pipeline.
This is some demonstration in simulation.
Um finally in the in the you know maybe the last four five uh minutes I want to talk about what I want to work on in the future. Um so I want to work on robotic applications that has real impact on people's everyday life on the real life constraints. So I think data problem is one major real life constraints we have.
Um, and I I'm always interested seeing robots, my algorithms or models deployed on the real hardware. Um, preferably on structure environments and I think I one of my main future focuses will still be robotic man manipulation and I believe robotic manipulation is uh the pearl of robotics. It's one of the most prized and challenging domain in robotics and I I think in order for us to achieve the goal of like generalist robots that can manipulate um that can do dextroous manipulation local manipulation as humans can do there are uh we need to do works on multiple fronts and one of the front I think is we need really scalable data collection pipeline for multimodel data collection. Um so last summer I was working at BDI now called Ray uh on just developing handheld device based on Yumi adding a tactile sensor and it can collect basically visual tactile demonstrations.
And then with this visual tactile demonstrations, what we can do is to replicate this task on the real robot.
And we show that the added tactile signals help it with task like tasking exertion because when you close the gripper, it's really hard for the camera to see what's going on under the gripper.
And fast forward to today, uh, after I joined Ry about two months ago, uh, we made a very impressive progress since last summer, we are developing the state-of-the-art handheld data collection device. Um, and we're also collaborating with people like uh outside this company to to scale up our data collection. So I think this will be an important and interesting area to to keep exploring. And when I say scalable data collection pipeline, I don't necessarily mean uh handheld data collection. I think T operation is also a variable uh interesting choice and both of them have their pros and cons.
And another front is I am a huge believer in simulation and I believe simulation and synthetic data is one of the key to really address the data problem in robotics. Um like I think right now we still have a lot of um challenge in simulating for example raw tactile signals or um deformable objects. If we can really simulate raw tactile signals I think that can enable a wide range of dextrous uh behavior.
Another thing is uh multimodal foundation models. Um I believe utilizing internet data by utilizing large pre-trained models on uh language or videos are uh will be very important to achieve generalist robots.
Uh to to to the best of all my knowledge I don't think right now we are at the point where we truly utilized the the common sense or the the physics that are learned by the large language models. most of the time we are still using language as like um uh a indicator of which task to perform. So I think uh there will be a lot of opportunities to work on this domain. And finally I think very importantly we need a very robust and comprehensive yet easy to spin up evaluation benchmark. just some simple codes that every lab can just spin up and then associate your model or your algorithm with a number and then this number can be compared uh across different labs. uh I think that will be very useful and this is something we need uh to really understand the model we learn to evaluate the model we train and with that I think um I this is the end of the presentation and I want to thank my amazing collaborators and thank all the audience for for taking your time to to attend this talk and I'm open for to questions Thank you. Thank you. Questions.
>> Uh, sure. Um, in the rehabilitation setting, have you thought about, it seems like a very natural goal to have continual adaptation. Have you thought about how to extend that kind of fshot learning into the continual adaptation setup?
>> Uh, like something along the line of sized learning like utilizing unlabelled data, >> something like that. But just in the sense that it it would be desirable if the device continues to adapt to the person's >> um signals in situ, right? So how have you thought about how we might approach that?
>> Yeah. Yeah, definitely. So um one I I mean like the the semis learning work is along that line. uh we also tried to um develop some unsupervised learning uh or a feedback mechanism for the patient. So basically one of the ongoing work we have is we'll provide a bottom for the for the patient to provide feedback. So after the model is being deployed the the patient can use the button to to inter to basically interrupt the model and then provides feedback. For example, if the model forgets to open and then the patients can say okay press the button and indicates now um when you see this data you should open your hand. So this is something we are we are studying uh right now and button is basically a very simple and binary uh feedback mechanism.
We can also use language to to give more uh uh fine grained uh feedback on how to uh adapt online the the model.
>> Does that make sense?
>> Yes. Thank you.
>> Thank you.
>> One question regarding the benchmarking you mentioned at the end when you say every lab should be easy easy easily able to do that. Um >> I always see the risk with benchmarks that if they become too specific people over at the bottom >> people over they starting they start overfitting.
>> Exactly. So people will >> overfit to your benchmarks and present better and better numbers but mainly because the benchmark doesn't capture enough variability right how do you think we can ever overcome that especially in the real world? Yeah. Um I think it's really hard. This is a very good question. I think we cannot avoid completely people just overfeeding to the benchmark. But I think even though people do overfeit to this benchmark, it does provide value uh even though it it overfits. So if you can overfeit to a specific benchmark, I think there are definitely things you learned that are useful for you to overfeit to that benchmark. And also I think this benchmark should also be continuously maintained. I mean when people are at the stage where they completely nail this benchmark it is probably the time to come up with a better benchmark which incorporates uh more challenging task by observing how people uh how how their model fail.
Um >> yeah.
>> Do do you see any potential for like using simulation? Because maybe in simulation if you get the realism done which is a big if as we know but if so then you can do much more automatic randomization and things like that and make the overfitting to the tasks harder.
>> I see.
I agree with you. I see huge potential in simulation but I I just I'm worried that right now given what people can do in simulation is not um challenging enough >> for as a benchmark. So I think it will eventually be a benchmark after we make it much better.
>> Yeah. No, that's a good point. Yeah, it was still not quite there yet.
>> Yeah, good question. Any more questions?
If >> nobody else, one question I could ask that's a bit more specific in the I think it was a tandem 3D where you had these objects and then the touch points, the contact points on them. It feels like an and if you put these objects onto that white pedestal, then actually the global position of where you make contact tells you already a lot of information about for example the size of the object and things like that which makes actually the classification much easier than if you would only do it based on shape. Did you >> Mhm.
>> Did you prevent it from doing that or did you give it the global contact point? Do do you understand what I mean?
>> Could you elaborate a bit more?
>> Okay. So, for example, let's say you have two two spheres.
>> Yeah.
>> And they one is like whatever that big and the other one is twice as big. And by putting the sphere >> on the pedestal.
>> Yeah.
>> You can already identify which of the two it is simply by the first touch based on how far away it is from the pedestal.
>> Yes, that's true. If you only the location is the same for both but this is not true for this case because our the location is actually randomized. So >> okay so the even the location of the this white pedestal on which the object is placed is also randomized.
uh that is not randomized but uh because it's it's connected to the to the the board but uh >> the object because it's object is using welcome to to >> to to attach to the the cylinder it location can vary and the orientation can also vary. So technically it can it can only occupy half of the um the the cylinder but >> okay >> so but inside inside our training simulation there's no like this cylinder stand and the the object can basically freedom move inside the workspace as long as you can get a first contact.
>> Mhm. you cannot like put it away such that the robot moves along the diagonal of the workspace and couldn't get a contact then that's an invalid episode >> but uh we we does randomize the location but during testing we have to uh like attach this thing to the workspace so it's limited testing could you imagine having a network that takes as input also object models like right now I guess your network is trained on this specific specific set of objects, right? But making it so that >> you say like, okay, you have five object models. Which of those is it?
>> Yes. Um, so that's an interesting point.
I would argue that our model is model based because during training we are basically keep loading this object in our RL training. So we are actually using the model. So, but I agree like if we want to maybe uh extend this pipeline to not just instance level identification to class level identification, we do need somehow to in in input more information to the policy >> because right now like we only have a limited am number of objects. What we want to do in the future is for example if you have different type of sofas or different type of bottles can you um identify on a class level right >> I think when we reach that level of difficulty then we do need to utilize models more explicitly >> yeah that makes sense >> I had another question about this work actually um >> I was wondering about the the strict necessity of the classifier module because if the explorer module has enough information to do efficient exploration, >> shouldn't it also have enough to do discrimination?
>> That's a very good question. This is actually one of our baseline. So one of our baseline we just give everything to the explorer. So um the explorer has to lead the exploration but also like have to determine when to stop the exploration and what is the output what is the identity of the target object. So it takes much longer to train uh because it doesn't uh it's not efficient as you separating you have a separate model of discriminator we train it specifically with a a label data set right um so yes that's a good point but it doesn't work as well as separating these two components >> okay thank you >> is there some connection you could make to like ectoritic kind of models in reinforcement learning in general Yeah, I think the idea of this is actually somehow from the actor critic setup.
>> Um, yeah, it's similar in concept.
>> Okay, thank you very much. I think we're right on time.
>> All right, let me let me stop sharing.
>> Thank you.
>> Awesome. Thank you everyone for joining.
>> Thank you everyone for your time. Bye.
>> Thank you. I'll be in touch.
>> All right. Thank you. Bye.
>> Have a good weekend.
>> You too. Bye.
Related Videos

Setting up a curved screen with Immersive Calibration Pro 4 and multiple cameras (P3D v4)
FlyerOneZero
23K views•2019-07-21

Jorge Mendez-Mendez: Unlocking Lifelong Robot Learning With Modularity (2023-10-05)
umassmlfl
237 views•2024-01-06

Northwestern’s MS in Robotics: Student Robotics Projects, 2023
NorthwesternEngineering
1K views•2024-05-31

"Perfect" Turns: Turning by the Gyro - FIRST LEGO League (FLL) SPIKE Prime + EV3 RePlay Programming
ZacharyTrautwein
94K views•2020-10-02

Gorkem Secer: TSLIP-based Deadbeat Running Control of Bipedal Robot ATRIAS
DynamicWalking-wv6qm
298 views•2018-06-22

Self-Driving Cars Need Lessons On Human Drivers | Maddie About Science
skunkbear
26K views•2018-08-21

Milrem Robotics’ THeMIS UGVs used in a live-fire manned-unmanned teaming exercise
MilremRobotics
99K views•2021-05-20

Robotic Guide Dog: Leading a Human with Leash-Guided Hybrid Physical Interaction
hybridrobotics3641
799 views•2021-07-29
Trending

YouTube Disabled Our Comments Again (Are Any Humans Left at YouTube?)
SpecialBooksbySpecialKids
39K views•2026-07-21

THE POLYGAMIST WAS ALMOST TOO MUCH MESS FOR ME | KennieJD
KennieJD
45K views•2026-07-21

YouTube Premium Isn’t Actually Ad Free
LMGClipsYT
41K views•2026-07-21

The REAL History Behind The Odyssey Will BLOW Your Mind! It's NOT a Myth!
metatronyt
20K views•2026-07-21