flashtensors

flashtensors

Hotswap large models on a single GPU

Description

Serving several models on one GPU means slow disk loads every switch, delaying the first token. flashtensors is a blazing-fast inference engine loading models from SSD to VRAM up to 10x faster.

It hotswaps a hundred large models on one GPU with minimal time-to-first-token impact, built on ServerlessLLM code.

Features



Fast loads:Up to 10x.

Hotswap:100 models.

Low latency:Quick first token.

One GPU:Saves hardware.
Tags:llmgpu