A system that automatically generates training data for humanoid robots by reconstructing real-world scenes using 3D Gaussian Splatting, sampling task-relevant waypoints, synthesizing robot motions with conditional diffusion models, and rendering egocentric views with visual domain randomization, enabling zero-shot transfer of synthetic-trained policies to real-world navigation and manipulation tasks.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
VLK: Learning Humanoid Loco-Manipulation from Synthetic Interactions in Reconstructed Scenes
Added:Welcome. We introduce VLK, a system that teaches a humanoid robot to see, listen, and act in the real world. Here is what it can do. Our first contribution is a pipeline that generates training data automatically inside reconstructed real world scenes. Step one, scene reconstruction. We scan a real room with an iPhone and reconstruct it using 3D gshian splatting. The result is a metric scale photorealistic copy of the environment. Step two, waypoint generation. Using the scene geometry, we sample task relevant waypoints. Feasible spots for the robot to walk to, pick from, or place onto. Step three, robot motion synthesis. We synthesize Unitry G1 whole body motions that follow these way points. For both navigation and box interaction using a conditional diffusion model. Step four, egocentric rendering. We replay the synthesized motions inside the 3DGS scene and render what the robot would see through its head camera. We also apply visual domain randomization while rendering, varying lighting, color saturation, and camera parameters. On the left, no randomization. On the right, with it, this helps the policy transfer from rendered images to real ones. Putting it all together, this pipeline generates 48,000 paired trajectories per scene fully automatically. Now, let's look at what those trajectories look like.
With paired data, we train a vision language kinematics policy. It maps what the robot sees and what it's told to do into the motion it should produce. The policy takes three inputs. The egocentric image, the language instruction, and the robot's current kinematic state, including wrist contact. During training, we use ground truth contact from the data set. At deployment, the contact is predicted auto reggressively from the model's previous output. We initialize from a pre-trained PI0.5 vision language action model and fine-tune it on our synthetic data. The policy predicts a 1second future whole body trajectory at 30 hertz joint angles, root pose, and wrist contact labels. A separate whole body tracker trained in simulation converts these kinematic targets into joint level commands the physical robot can execute.
The whole pipeline runs in real time with 31 millisecond inference latency.
That completes the policy. After training, we deploy our model in the simulation and real world.
finally we show that the VLK policy trained on purely synthetic data can be executed in the real world zeros task one navigation given a language instruction the robot walks to a target object Even with layouts it has never seen during data generation.
On the left, the robot starts from a random pose. On the right, the language instruction was never seen during training. Both work. Zero shot.
Task two, box lifting. The robot picks up boxes of three different sizes from the floor, small, medium, and large, and puts them back.
Task three, multi-stage tasks. By streaming language instructions in real time, we chain atomic skills, walk, pick, carry, place into longer behaviors.
Task four, robustness to disturbance. On the left, we change the scene layout while the robot is acting. On the right, the robot operates under flickering disco light.
Task five, long horizon tasks. By chaining navigation and pick and place with real-time language, the system completes multi-minute behaviors.
VLK shows that synthetic interactions in reconstructed scenes can supervise real humanoid loco manipulation. Thanks for watching.
Related Videos

Setting up a curved screen with Immersive Calibration Pro 4 and multiple cameras (P3D v4)
FlyerOneZero
23K views•2019-07-21

Robot Learning with Sparsity and Scarcity
allenai
379 views•2025-10-14

Jorge Mendez-Mendez: Unlocking Lifelong Robot Learning With Modularity (2023-10-05)
umassmlfl
237 views•2024-01-06

Northwestern’s MS in Robotics: Student Robotics Projects, 2023
NorthwesternEngineering
1K views•2024-05-31

"Perfect" Turns: Turning by the Gyro - FIRST LEGO League (FLL) SPIKE Prime + EV3 RePlay Programming
ZacharyTrautwein
94K views•2020-10-02

Gorkem Secer: TSLIP-based Deadbeat Running Control of Bipedal Robot ATRIAS
DynamicWalking-wv6qm
298 views•2018-06-22

Self-Driving Cars Need Lessons On Human Drivers | Maddie About Science
skunkbear
26K views•2018-08-21

Milrem Robotics’ THeMIS UGVs used in a live-fire manned-unmanned teaming exercise
MilremRobotics
99K views•2021-05-20
Trending

2.4 BILLION Records Got Leaked...
DeepHumor
15K views•2026-07-22

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Should I buy a Sawmill?
essentialcraftsman
29K views•2026-07-22

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23