
Description
Most robot policies map pixels straight to actions without understanding what happens next, so they fail on new tasks. LingBot-VA is a causal video-action world model that predicts future frames and robot actions together for generalist control.
It generalizes across manipulation tasks, published at RSS 2026.
Joint prediction:Frames and actions.
Causal:Understands consequences.
Generalist:Many tasks.
It generalizes across manipulation tasks, published at RSS 2026.
Features
Joint prediction:Frames and actions.
Causal:Understands consequences.
Generalist:Many tasks.
