Back to library

Robotics / NVIDIA

DreamDojoA Generalist Robot World Model from Large-Scale Human Videos

Three key questions about this paper

What problem does DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos address?

Shenyuan Gao et al. · ICML 2026 Spotlight · arXiv:2602.06949

What evidence supports the main claim in DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos?

The strongest generalization test asks whether generated futures look more physically correct and follow actions better in scenes outside the robot-training distribution. Across 50 samples judged by 12 volunteers, DreamDojo-14B beat the Cosmos-Predict2.5 baseline in 73.50% of physics comparisons and 72.55% of action-following comparisons. DreamDojo-14B also beat the 2B version 72.50% and 65.53% of the time on those two criteria. [Paper §4.4, Table 4]

What limitation should readers know about DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos?

The boundary matters: two “novel” evaluation sets used image-edited backgrounds, the preference study was small, and video plausibility is not the same as successful physical execution. The paper demonstrates broader simulation within six designed benchmarks and a few downstream robot settings; it does not establish universal physics, guaranteed long-horizon accuracy, or safe zero-shot deployment.

2 new free reports left todaySubscribe to Pro for unlimited reading and 10 new paper explanations each month.Upgrade to Pro