
Description
Robots that learn manipulation fail when the object or scene changes. DiT4DiT is a vision-action model combining a video generation model with flow-matching action prediction, imagining how the scene evolves before choosing actions.
Physical common sense from the video model helps it generalize to new objects and scenes; code is open.
Video generation:Predicts future frames.
Flow matching:Smooth actions.
Generalization:New objects and scenes.
Physical common sense from the video model helps it generalize to new objects and scenes; code is open.
Features
Video generation:Predicts future frames.
Flow matching:Smooth actions.
Generalization:New objects and scenes.
