LingBot-Vision

LingBot-Vision

Self-supervised vision backbones for spatial perception

Description

General vision backbones often fall short on detail and geometry for dense spatial tasks like depth and segmentation. LingBot-Vision is a family of self-supervised ViT backbones for dense spatial perception, from ViT-S/16 up to 1.1B parameters.

The pretrained backbones serve downstream robot perception and 3D understanding.

Features



Self-supervised:No labels needed.

Sizes:ViT-S to 1.1B.

Dense tasks:Depth, segmentation and more.