WAM-TTT offers a sophisticated yet practical solution to the generalization problem by enabling robots to adapt to new environments through real-time updates from human videos. This shift from static models to dynamic test-time training marks a significant leap toward truly scalable embodied intelligence.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
WAM-TTT: Teaching Robots to Adapt to New Environments by Watching Humans | GALBOT AstraBrain
Added:Galbot introduces Astro Brains' latest breakthrough in scalable robot deployment, VOMTTT world action model test time training.
The world's first test time training paradigm without robot action data [music] at deployment, breaking the bottleneck of efficient, scalable, cross-environment deployment. Simply record human action videos, no annotations required. Robots can now learn new skills by just watching demonstrations in the target environment. Training embodied foundation models consists of two stages, pre-training and post-training.
Post-training today heavily [music] relies on data collected in data training centers.
For the same skill, when a trained robot is deployed into a new environment, performance often suffers from a sharp drop.
Traditional teleoperation-based deployment solutions are expensive, time-consuming, and require specialized equipment. Moreover, the resulting skills often have limited generalization. To address these challenges, we introduce VOMTTT, a revolutionary new framework. We enable any users to teach robots by simply recording human videos. Human videos are low-cost, easy to collect, and ready for immediate adaptation. By simply watching unlabeled human demo videos, robots can adapt to new environments without robot action data or manual annotation. In our data scaling ablation study, with only 100 robot trajectories combined with 100 human videos, the average task completion score [music] reaches 74.1%.
Performance not only remains comparable, but even improves [music] over using robot data alone. This means that videos can replace expensive robot teleoperation data at a 1:1 ratio.
Furthermore, we significantly reduce dependence on human data annotation.
Existing approaches are often limited by high-quality action labels and precise motion retargeting.
VOMTTT breaks through this bottleneck.
It learns directly from raw human videos, the simplest and most effective method.
When adapting to new environments, traditional supervised fine-tuning SFT faces a high risk of catastrophic forgetting. VOMPTT introduces a new paradigm. Foundation model backbone remains frozen plus lightweight fast weight adaptation. The pre-trained model remains completely frozen. Only lightweight fast weights are updated.
Compared with SFT, this approach preserves the generalization ability of the pre-trained model. This decoupled design not only prevents forgetting, but also stores newly acquired [music] skills as reusable memories within the model, laying the foundation for scalable deployment of general-purpose robots.
In real-world cross-environment evaluations under shifts in lighting, workspace, and object instances, the baseline approach that directly uses human videos as contextual input suffers catastrophic [music] generalization failure with only 14.7% performance retention.
In contrast, VOMPTT demonstrates overwhelming advantages, achieving 75.6% performance retention. These results systematically prove that learning through adapting the fast weights provides substantially stronger generalization than relying solely on in-context learning.
Easy data collection, one-click adaptation, fast deployment. Galbot is pioneering a revolutionary new approach to enable the scalable deployment of embodied intelligence.
Related Videos

Setting up a curved screen with Immersive Calibration Pro 4 and multiple cameras (P3D v4)
FlyerOneZero
23K views•2019-07-21

Robot Learning with Sparsity and Scarcity
allenai
379 views•2025-10-14

Jorge Mendez-Mendez: Unlocking Lifelong Robot Learning With Modularity (2023-10-05)
umassmlfl
237 views•2024-01-06

Northwestern’s MS in Robotics: Student Robotics Projects, 2023
NorthwesternEngineering
1K views•2024-05-31

"Perfect" Turns: Turning by the Gyro - FIRST LEGO League (FLL) SPIKE Prime + EV3 RePlay Programming
ZacharyTrautwein
94K views•2020-10-02

Gorkem Secer: TSLIP-based Deadbeat Running Control of Bipedal Robot ATRIAS
DynamicWalking-wv6qm
298 views•2018-06-22

Self-Driving Cars Need Lessons On Human Drivers | Maddie About Science
skunkbear
26K views•2018-08-21

Milrem Robotics’ THeMIS UGVs used in a live-fire manned-unmanned teaming exercise
MilremRobotics
99K views•2021-05-20
Trending

NOLAN WELLS AUTOPSY: DEEP TISSUE DISCOLORATION BACK OF HEAD, NECK BONE MISSING
nancygrace
561K views•2026-07-22

Big Tech's Biggest Gamble Is Finally Falling Apart
houseofel-ai
79K views•2026-07-22

we're almost finished the house (ep.125)
JennaPhipps
347K views•2026-07-22

We flooded a field - the results blew our minds
MossyEarth
73K views•2026-07-22