
Description
Humanoids that walk to a table and reach for something usually use separate systems for walking and manipulation, joined awkwardly. WholebodyVLA proposes a unified latent vision-language-action model controlling whole-body locomotion and manipulation.
It learns from action-free videos, published at ICLR 2026.
Whole body:Walk and manipulate.
Unified:One policy.
Video learning:Unlabeled data.
It learns from action-free videos, published at ICLR 2026.
Features
Whole body:Walk and manipulate.
Unified:One policy.
Video learning:Unlabeled data.
Screenshots
Tags:robotics
