
Description
When multimodal LLMs drive, reasoning in text about what the car ahead will do loses spatial detail. FutureSightDrive has the model imagine the future frame first and plan from that image, replacing text-only reasoning with a visual spatio-temporal chain of thought.
From Xi'an Jiaotong University, a NeurIPS 2025 Spotlight with training and inference code.
Visual CoT:Imagine the future.
Spatio-temporal:Spatial detail kept.
Planning:Trajectories out.
From Xi'an Jiaotong University, a NeurIPS 2025 Spotlight with training and inference code.
Features
Visual CoT:Imagine the future.
Spatio-temporal:Spatial detail kept.
Planning:Trajectories out.
