
Description
For a robot to answer where that cup on the table went, it must understand objects and remember when and where it saw them. DAAAM uses SAM segmentation, tracking and vision-language models to build 4D dynamic scene graphs in real time, with queryable descriptions for every object.
It leads on NaVQA and SG3D, has a ROS 2 interface and appears at CVPR 2026.
Live descriptions:Objects as seen.
4D graphs:Space and time.
Tracking:SAM and BotSort.
ROS 2:Robot-ready.
It leads on NaVQA and SG3D, has a ROS 2 interface and appears at CVPR 2026.
Features
Live descriptions:Objects as seen.
4D graphs:Space and time.
Tracking:SAM and BotSort.
ROS 2:Robot-ready.
