TurboQuant-GPU

TurboQuant-GPU

5x KV cache compression for LLM inference

Description

Long-context LLM inference fills VRAM with KV cache, limiting length and concurrency. TurboQuant-GPU delivers 5.02x KV cache compression for LLM inference on any NVIDIA GPU.

It quantizes using random orthogonal rotations and more, with cuTile kernels, automatic PyTorch fallback and a PyPI package.

Features



5x compression:Much smaller KV cache.

Any NVIDIA GPU:Broad support.

Fallback:PyTorch path.

PyPI:pip install.
Tags:llm