Install our extension to search inside any video instantly.

What we're working on - Menlo Community Talk

Added:
113 views6likes46:37menloresearchOriginal Release: 2026-07-10

A robot-first multimodal encoder that fuses vision, proprioception, audio, and force sensors into a shared latent representation enables robots to reason about tasks more effectively than vision-only encoders, as demonstrated by experiments showing that vision-only latents cannot recover force information that cameras cannot directly observe, and that a single embodiment-agnostic encoder trained across multiple robot types outperforms specialist encoders.