
Description
Video world models predict 2D frames, while robots need to know how objects move in 3D. PointWorld represents the world as 3D point clouds and predicts how every point moves under robot actions.
Trained at scale, it generalizes to in-the-wild manipulation for planning and policy evaluation.
3D prediction:Per-point motion.
Scale:Real-world generalization.
Planning:Best actions.
Trained at scale, it generalizes to in-the-wild manipulation for planning and policy evaluation.
Features
3D prediction:Per-point motion.
Scale:Real-world generalization.
Planning:Best actions.
