WholebodyVLA

WholebodyVLA

Unified latent VLA for humanoid loco-manipulation

Description

Humanoids that walk to a table and reach for something usually use separate systems for walking and manipulation, joined awkwardly. WholebodyVLA proposes a unified latent vision-language-action model controlling whole-body locomotion and manipulation.

It learns from action-free videos, published at ICLR 2026.

Features



Whole body:Walk and manipulate.

Unified:One policy.

Video learning:Unlabeled data.