
Description
General vision backbones often fall short on detail and geometry for dense spatial tasks like depth and segmentation. LingBot-Vision is a family of self-supervised ViT backbones for dense spatial perception, from ViT-S/16 up to 1.1B parameters.
The pretrained backbones serve downstream robot perception and 3D understanding.
Self-supervised:No labels needed.
Sizes:ViT-S to 1.1B.
Dense tasks:Depth, segmentation and more.
The pretrained backbones serve downstream robot perception and 3D understanding.
Features
Self-supervised:No labels needed.
Sizes:ViT-S to 1.1B.
Dense tasks:Depth, segmentation and more.
Tags:computer-vision

