DiT4DiT

DiT4DiT

Vision-action model for robot manipulation

Description

Robots that learn manipulation fail when the object or scene changes. DiT4DiT is a vision-action model combining a video generation model with flow-matching action prediction, imagining how the scene evolves before choosing actions.

Physical common sense from the video model helps it generalize to new objects and scenes; code is open.

Features



Video generation:Predicts future frames.

Flow matching:Smooth actions.

Generalization:New objects and scenes.