Humanoid robots require a comprehensive perception stack to operate autonomously in unstructured environments alongside humans, consisting of three core capabilities: (1) Visual SLAM for precise localization and mapping with sub-5cm accuracy, (2) Obstacle detection and free space understanding to ensure safe navigation, and (3) Scene understanding with AI features to identify objects, people, and predict their intentions. This stack enables robots to answer three critical questions in real-time: 'Where am I?', 'Where can I move safely?', and 'What's around me?' The stack relies on multimodal sensor fusion combining cameras, structure light, IMUs, and joint encoders, with simulation-to-real transfer using physically accurate sensor models to accelerate deployment from months to days.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Alon Zehngut - Building the Perception Stack for Humanoid Robots
Added:director of product robotics at real sense will present the visual cortex for physical AI.
This session will explore why depth perception, spatial intelligence, and GPU accelerated perception software are essential for humanoids that need to navigate safely and predictably through warehouses, factories, and shared workspaces.
Please welcome Alan Zer.
>> Yes.
Thank you very much.
>> Can you hear me, guys?
>> Yeah, >> it's working.
Well, they told me the Thursday is the last day. It's going to be quite empty.
They were right. But thank you. I appreciate for you coming at all. See this one. Are you doing robotics?
Humanoids, drones, humanoids.
Cool. Awesome.
So, there's a quite small audience. So, um feel free to step in, ask question during the show. I won't do like a straight on presentation and then if Q&A can be more interactive, feel free to engage. Um keep it short, 15 20 minutes.
You can ask questions. So, this session is going to be building the perception st for humanoid robots. This is what we do. And maybe just a short introduction about myself. I'm a product director for Real Sense. I'm sure some of you heard about Real Sense. Who knows Real Sense?
>> Yeah, it's good. So, I'm doing the go to market and strategy for the robotics portfolio, physical AI, and I've been in industry for 15 years doing um AMRs, mobile robots, and also unmanned uh vehicles in the different industry. So that's UAV, drones, unmanned ships, stuff like that. And the last six, seven years I spent in Argo Robotics doing a visual slam mostly for mobile robots.
And this is my specialty. Um how we can handle dynamic crazy environments to be very um robust and accurate versus 2D light sometime get lost in those repetitive eyes or very dynamic environments. So that's about about visual slum and uh I do believe that humanoid will be somewhere around consumer in the next three years. It might be too ambitious but this is what I'm thinking is going to be uh soon. And just about real so I'm sure you're familiar with the cameras. Some of the companies here are using real sense from the small medium range to the close proximity on the wrist of the humanoids. So we do cameras but what we do beside cameras is also visual slum perception. That's one I talk about. We're doing a lot of work in activity and safety. I put some words about it as well. Biometrics. We are positioning in um airports. We're doing face authentication solutions which is also going to be very relevant to humanoids very soon. Uh of course within video we do simulation and open SDK and a lot of stuff. And it's important to understand we're doing from the camera ID. So the metal sheet, the lenses, the ASIC or the system on chip, but we also hold the whole ST. So the algorithm that runs on the camera, the AI features, the perception features that run on the camera. And we now introducing to the market more and more software solution STS for the full solution like autonomy tech. Any questions? Awesome.
So I'm sure you all know about agility robotics. They just announced going public uh yesterday. Good friends of ours. They're using like six wheels cameras. Pretty good example. And I think Agility Robotics is one of example I really like because they are the first one that actually brings value to the market. Instead of doing some shows and cool dancing, they are walking at GXO, Toyota facilities and stuff like this carrying almost 100,000 toes per day which is crazy, right? Think about human being doing this. It's quite a hard uh job.
But if you look at it in a high level and you scope this one, this use case is working in a very structured environment. So you saw those one, they're walking in a fence, right? It's very closed environment, very structured and the task they are doing is very repetitive. So the test same task every day, right? They're doing the toads from one place to another put it on AMR and this is the exactly the one specific task they were designed to do. So I don't say it's easy but it's that's the bar they are doing today >> right >> um and again people is outside those fences there is no really really um humanoid people collaboration so the humans working outside the fences and this is what we see today in actual real deployments just a short movie to um capture this one out so as you can see this is the digit and he cares and he goes from this convey to another place and you can see on the floor and you can see on the rails they're using QR codes they're using QR codes to have a really um short range navigation right they navigate in a relative position from one QR code to another means they don't usually um facilitate the whole database of the map of the environment because they don't need to they go from point A to point B and doing this repetitive task but what we are seeing today and what automate show trying to bring us is humanoid robots are leaving the cage.
They want to get the humanoid engage with people. They want to get humanoids doing more general purpose tasks. They want us to have the laundry set up. They want us to do the coffee right and this um photo I took not I took made by Jiminy sorry try to illustrate the mutual work. So humanoid is alongside human in a complex unstructured environment, right? They're going to walk together and it's going to be more general purpose task. That means not only toads carrying from one place to another. They're going to help you unscrew and unscrew the drivers. They're going to help you bring some equipment.
They're going to do things for you with you. And that requires high level intelligence, high level perception to be able to do so, right? Um and again human beings we are unpredictable just like robots.
So this is where we heading and I'm sure you saw the 1x figure um video and let's be honest right this is teleop this is not autonomous mission but it's a preview to the future what the humanoids will do in the future right they'll probably be in the house do our laundry do general purpose but this require a high level intelligence and perception because they need to understand where they are exactly in the space they need to interact safely with human beings And this is really really much more complicated than the task they are doing today. So this is where we headed in the future. And what I'm here today to tell you is how we can achieve this high level of perception and how we building the perception stack for the humanized robot. So in order to do this autonomous task, the robots have to answer three main questions in real time and all the time. And the first question will be where am I? Right? I need to know exactly during the whole mission where am I in the space where am I in the room where am I in the office in a very robust accurate way non-drift position.
So every time I go to a specific point I get a precision of 5 cm exactly because if I get my position in a 10 cmters off my task will be failed probably. And how we achieve this high position using visa visual simultaneous localization and mapping these six degrees of freedom six xyz position and the orientation pitch a roio and this uh this is the answer to the question where am I second question will be where can I move and I added recently the word safely not sure if you all see invidia just published with agility the halls OS everybody talks about safety How can I do this operation in a safety manner? That means when I work with human, if my batteries run out, I fell down, I don't injured, I don't hurt anyone. And there's a lot of evolving safety standout. Not sure if you're well aware, there's AI functional system and AI standout that will be introduced in the next year. A lot of people working on that one. We are actively in this domain. How can a human being interact with a humanoid in a safe manner outside a cage in a collaborative job? And to achieve this one we have to understand what is the free space and I'm sure you're familiar with drill sense does obstacle avoidance occupancy grid free space map but also we now introduce um AI features such as people detection I need to understand if this is obstacle is a static one or this actually the mission a person I need to engage with.
So this is very important to understand what I'm doing in travel my mission. And the third question what's around me it's more in the high level scene understanding uh layer it's not only there's an obstacle if it's a human or not is what is the intentions this human being is going to be static and stand here it's going to be walk through as he trying to put this object to the other one and this will give me all the high level understanding that I need uh to do so and let's talk about human uh humanoid localization. I've been doing visual slam for AMRs for the last six years and I must tell you it's has been achieved. There are visual slam deployments with real AMRs around the world and AMR are much more easy versus humanoids. If you look at the encoders, the odometry, they are very clean, accurate. They navigate on a flat surfaces, very predictive. The motion speed is very very um predictive and accurate. And the humanoids just break the equation. They work with a noisy joint encoders. Their kinematics is much more noisy. The mission requires much more flexibility. Climbing the stairs going uneven surfaces. That's crazy. And how we handle it now is introducing VIO which is stand for visual inertial autoometry is fusing the vision data together with IMU and the joint encoders to have an more accurate and robust pose estimation. And this is only one fundamental building blocks that goes into the vislam pipeline. So I'll show you in a minute how we take kislam from Nvidia. This is the vio building blocks and we populate it into the real vislam to have the bestin-class visual slam localization mapping and autonomy tech.
So what you can see in the video here is quite very interesting. On the left you see the humanoid of Limix Dynamics, one of our partners we did collaborate with walking down the office and you see the colorful pictures is the death point cloud or the recent camera and you can see the map is building in real time and on the center all these colorful brilliant beautiful dots. These are the features. So the features the camera is tracking and how visual vio work is for frame to frame. We look at the difference between what I saw in frame number one, frame number two, frame number three. And this is how I can estimate my post trajectory, right? In a three-dimensional fuse it with an IMU. We have IMU Bosch IMU on each one of the cameras. And again, we have the joining quarters from the humanoids. You can get a multimodality fusion of pose estimation.
And this is only the basic layer to understand. And using only vio by the way we can estimate in less than 1% path travel the pose accuracy which is quite crazy but now fuse it with a vis and create a database you can get a global position accuracy of sub 5 cmters during the whole mission 247 during the week which this is what we're heading to and then again for the obstacle detection if I want to know objects this again show obstacle Oh, this is people detection. Now the robot's not can only see obstacles. It can understand whether it's a human being, whether it's a static obstacle and how to navigate and he can plan his uh mission accordingly.
So that's was capturing last week with the D5.
Let me show you this again. I want to show this again because I want you to look at the field of view. This is 87 90 horizontal field of view. You're familiar with the Real Sense camera. And some of you that follows Real Sense probably know we just introduced our new camera, the D585, which has a 120 horizontal field of view. And you can see now this movie.
Look how wide it is. And the depth go all the way from 15 cm to 10 meters. And this is significantly improve the performance point and the ability of the robot to see much further away, much wider away, closer away and to identify in real time where there is an obstacle, people and what is the accurate um distance for each one of those one. And this is how you achieve reliable safe uh perception for autonomous navigation.
And again, I'm almost close to finish. I want to show you a video we did with Limix dynamics. So they were using 436 global shutter uh medium range real sense cameras together with Nvidia AGX compute integrating the VIO. So we creating a map of the office right the robot was mapping the office creating its own database of the vlam and then it was able to navigate and this is the first movie it's not teleop it's fully autonomous tech there's no tele operation no joysticks nothing fully autonomous mission of a robot navigates in the office aware about the obstacle knows that this is a people aware about the situation aware about it mission this is how we achieve um the full autonomous back. He knows now this is a static object versus the people. This is how now he plans better its route and um knowing its position in accurate matter, he knows that it was passing this aisle on this lobby to be better optimized uh towards its mission. I'll let this run another 20 seconds. Any question about this one?
Okay, we can leave it to the end.
This is the RGB feed. This is the depth.
This is the um robots do a full autonomous mission.
Then we cover the ST and what we see now in the world that you cannot do everything by yourself. You cannot take the robots and do hours of training, hours of mission. And this is where the simulation kicks in. And what we see that Nvidia introduced the Omniverse and SIM Isaac SIM get really really effective now to accelerate deployments instead of wasting and and participating the site itself and train the robot for years. You can do it with simulation and to get achieve really good simulation you have to use the real the real um inputs. So real sensor input which is not going to be synthetic. If you're going to simulate your robot with synthetic data, you train it. You train to go to the real world and it will break. Why? Because you didn't part you didn't accept anticipate the noise and the real world characteristics. So what happened? You do a lot of synthetic, you go to the warehouse and all of a sudden the robot behaves different what it in simulation. So what we were working on the last six months is how you can simulate the real sense cameras in the omniverse in a really really accurate and the physics are very real. So we actually introduce the noise level and the physics of the real camera. This is example this is Nvidia movies right you can see on the left the physical model and the digital twin. So they train this industrial arm to do a specific task, right? They do it for a simulation, it for the real before we introduce the noise level. If you move the piece a little bit, 5 m 5 degrees angle, it will fail because now it looks different in the real world versus the simulation.
But if you look at this movie, this is a completely synthetic scene mimicating the exact output you will get from a real sense D55 camera. Right? You can see the actual real noise, the waves you are familiar with in the long range of the depth camera. So in five meters, you know the depth is less accurate than in one meter.
And when you introduced a real physics and a real noise to this world, now you can do real simulation that shorten your deployments and get you um mass fetcher deployments much more reliable.
So let's just do a quick recap what we discussed. Humanoids are transitioned from unstructured environment to the real world unstructured space working alongside people. They need to be fully autonomous and to do so they have to have a vision-based reliable perception.
The full perception how we achieve it visual slam for localization navigation free space obstacle detection and the second third layer is going to be the scene understanding. So AI feature I want to understand if it's a paper it's a static object and I want to understand the trajectory and the velocity. This is what we build in the perception stack.
And again the sim to real is the physically accurate sensor model to actually compress the field testing into days versus uh months. Thank you very much for participating.
>> Thank you.
>> If you have any questions I'm here in the next five minutes.
>> Any questions?
That's a good question. Thank you. So, if you didn't visit the real booth, please go ahead in the no hall. It's 120 36 live demo. Great people to see. I'll probably stop by there.
It's been quite the level of sensors and sensors.
Is it just >> Thank you for the question. I think it's a good one. Um we're first on a really great belief of multimodality and sense of fusion. That means you don't want a single technology that be in charge of everything. Multimodality is the key.
When one technology fails, the other one compensates. So that that's critical, right? And when I say multimodality, we are fusing the camera, we are fusing structure light, so the IR projectors, we are fusing the IMUs on the cameras, we are fusing the join encoders. So that's a multimodality. But the key fundamental technology that I believe that will allow the autonomous operation is vision based because it give you much more texture much more understanding right even if you go to the LAR which was still the last two days the wooding technology it's quite reliable it's quite accurate it's walking a pitch death um in a very dark environment but it always see 2D just a planner right now you put a boxes in the warehouse the lighter don't see anything and again you go 3D lighter now you go see the ceiling you go see everything that's good but even the 3D light I don't understand whether this people is a person or this one is a obstacle and again even if they able to see now with a 4D imaging LAR and you can put an RGB sensor in it there is not enough resolution now to identify specific objects I want to be able to identify this is Chris with his crazy lobster hat and it's an authorized person to go into my booth versus all their competitors than otherwise. So that's only camera I can get right. So this is why I believe camera is the key fundamental and definitely multimodality with different sensors um are the key to success.
>> So two questions. So one is Are we seeing like a crossation here?
That's a great application.
Where do you see the best as far as you see in the office? I mean, the first thing I would do if I know What are the practical ROI for?
>> Thank you. There's two great question.
Let me start with the first one.
Automotive industry.
Um well, we see a lot of transition from automotive technology to the robotics and no doubt you know there's a camp between the vision camp Elon Musk Tesla and Mobilei. I'm from Israel. All the reasons R&D is from Israel. So this is the vision camp versus the lighter camp and definitely we see a lot of adoption from the automotive industry into the robotics industry that's go for gym connectivity came from automotive into robotics industrial we see a lot of standards the safety standard I mentioned so we started from the 2602 of the automotive and now it's populating into AI system into robotics so definitely there's a lot of adoption from the automotive industry into robotics from the computer level the technology level and the interface this level and now even to the regulation. So it's kind of a subset case study of the automotive industry. Each one have a different challenges of course. So that's definitely uh the case regarding ROI. I think well you read the news and you all read article papers. Now the humanoids are working in industrial manufacturing in manufacturer line and you want to see them more and more in healthcare. You want to see them in retail. This is really the ROI when you can get span out high volume. But if you ask the human manufacturer, they want to go consumer, right? The unit tree want to sell not just 5,000 bots per year, they want to sell million bots that everyone in his house. So depending how you measure the ROI um but I can see the ROI is working a very um repetitive in all the jobs that you and I don't want to do uh to make the people have more fun and more time.
a lot of good stuff.
Can you share more about simulation how sensors use for simulation building?
>> Sure, that's a good uh question. We see simulation as a key key um fundamental building blocks for um fast deployment and scale. And what we did in this domain is work really hard to take all the source code and all the deaf pipeline that runs on the camera into the Isaac sim. That means the Isaac sim is now um has the real characteristics of the lens of the camera of the physics. This is what they do, the Nvidia guys, right? They invest a lot of hours in getting the actual physical model of the lens, physical model of the chassis of the camera. And we took and populate the full real pipeline that works on the camera into Isaac scene to make sure when you do a simulation now, you can have much more close to the real um input and data.
Thank you for the very nice presentation. I'm from surgical robotics field. Uh I'm interested in the monitor the situation in operating room and steps.
uh what I'm curious about is technology but uh we are considering about >> what you think about between your technology >> is this go just make sure for static application volume monitoring space well I would say it's definitely depends on on the requirements and the product if you need a detection ranges more than 10 meters and you want to see in my whole room 50 m I want to detect objects far far away you go light up there's no doubt the stereo has limitation it doesn't see and detect very very good object in very good accuracy more than 12 meters let's go when I light up but If you want to get field of view, if you want to understand who is the people, what are they doing, if you want to get analytics, if you want to track those ones, if you want to get much more information out of this, definitely vision based technologies are um better. Um that's the key answer.
>> I'm a huge enjoy mention what advances what technological area [snorts] >> I would say it's two first one is safety and regulations till today there's limitation of deploy those ones outside the fence because people are afraid that they will get injured, they will get damaged and there's a lot of strict regulations that um blocks you of deploying that one. So there are still today all the robots that's been deployed or trained or monitored. They have a lot of restrictions. This is because the regulation restricts you to there's a lot of work that I'm aware of that people are doing to get safety enabled um systems in one year from now and the adoption will probably take another two years. So that's one. I think safety standards, safety components, safety products will be introduced and deployed the next three years. That's first and the second one is the um the understanding of where the robot is. So that's visual slum deployments. We're just talking with all the vendors this this uh automate show from agility to everyone else and they want vis now. They were deployed in um warehouse in a very restricted area.
They want to go wild. Once they introduce this down, they're going to do autonomous navigation outside. That's the second one that going to allow them to be fully autonomous. And there's a third one, all the VA in the training.
Now, people getting this robot much, much more smarter. And this is going to go so much faster than what we see in the last three years.
>> Thank you so much for >> Thank you guys.
Related Videos

Setting up a curved screen with Immersive Calibration Pro 4 and multiple cameras (P3D v4)
FlyerOneZero
23K views•2019-07-21

Robot Learning with Sparsity and Scarcity
allenai
379 views•2025-10-14

Jorge Mendez-Mendez: Unlocking Lifelong Robot Learning With Modularity (2023-10-05)
umassmlfl
237 views•2024-01-06

Northwestern’s MS in Robotics: Student Robotics Projects, 2023
NorthwesternEngineering
1K views•2024-05-31

"Perfect" Turns: Turning by the Gyro - FIRST LEGO League (FLL) SPIKE Prime + EV3 RePlay Programming
ZacharyTrautwein
94K views•2020-10-02

Gorkem Secer: TSLIP-based Deadbeat Running Control of Bipedal Robot ATRIAS
DynamicWalking-wv6qm
298 views•2018-06-22

Self-Driving Cars Need Lessons On Human Drivers | Maddie About Science
skunkbear
26K views•2018-08-21

Milrem Robotics’ THeMIS UGVs used in a live-fire manned-unmanned teaming exercise
MilremRobotics
99K views•2021-05-20
Trending

we're almost finished the house (ep.125)
JennaPhipps
347K views•2026-07-22

We Finally Know Where Saturn’s Rings Came From
astrumspace
79K views•2026-07-22

BIG BET: Cathie Wood goes ALL IN on Elon Musk
FoxBusiness
89K views•2026-07-22

MIC DROP: Smithsonian Director Called Out For Woke Propaganda
TheAmalaEkpunobi
37K views•2026-07-23