
Description
Video world models encode each frame as hundreds of tokens, exploding compute for long videos. DeltaTok from Amazon proposes delta tokens where a frame is worth one token, making generative world modeling far more efficient.
It encodes only changes between frames and supports depth and segmentation; a CVPR 2026 Highlight.
Delta tokens:One per frame.
Efficient:Long videos feasible.
Multi-task:Depth and segmentation.
It encodes only changes between frames and supports depth and segmentation; a CVPR 2026 Highlight.
Features
Delta tokens:One per frame.
Efficient:Long videos feasible.
Multi-task:Depth and segmentation.
