
Description
Tell a robot to put the cup in the second empty spot left of the plate and it usually can't parse the spatial relations. RoboRefer is a vision-language model for robots with spatial referring and reasoning that pinpoints the places and objects an instruction describes.
It combines depth and multi-step reasoning for complex spatial relations; published at NeurIPS 2025.
Spatial referring:Precise locations.
Reasoning:Multi-step.
Depth:3D awareness.
It combines depth and multi-step reasoning for complex spatial relations; published at NeurIPS 2025.
Features
Spatial referring:Precise locations.
Reasoning:Multi-step.
Depth:3D awareness.

