
Description
Robot data is slow and expensive to collect, while endless videos of humans doing things go unused. DreamDojo is NVIDIA's generalist robot world model learned from large-scale human videos that predicts future frames from actions.
Use it to evaluate policies, generate training data and plan; published at ICML 2026.
Human videos:Massive data source.
Action-conditioned:Predicts frames.
Policy evaluation:Test in imagination.
Use it to evaluate policies, generate training data and plan; published at ICML 2026.
Features
Human videos:Massive data source.
Action-conditioned:Predicts frames.
Policy evaluation:Test in imagination.
