
Description
Reconstructing a scene as video streams in is either slow or drifts over time with classic methods. LingBot-Map is a feed-forward 3D foundation model for streaming reconstruction using a geometric context transformer as frames arrive.
The paper was an ECCV 2026 oral and best paper candidate, with code and demos.
Streaming:Maps as video arrives.
Feed-forward:No per-scene optimization.
Geometric context:Consistent over time.
The paper was an ECCV 2026 oral and best paper candidate, with code and demos.
Features
Streaming:Maps as video arrives.
Feed-forward:No per-scene optimization.
Geometric context:Consistent over time.

