
Description
Long-context LLM inference fills VRAM with KV cache, limiting length and concurrency. TurboQuant-GPU delivers 5.02x KV cache compression for LLM inference on any NVIDIA GPU.
It quantizes using random orthogonal rotations and more, with cuTile kernels, automatic PyTorch fallback and a PyPI package.
5x compression:Much smaller KV cache.
Any NVIDIA GPU:Broad support.
Fallback:PyTorch path.
PyPI:pip install.
It quantizes using random orthogonal rotations and more, with cuTile kernels, automatic PyTorch fallback and a PyPI package.
Features
5x compression:Much smaller KV cache.
Any NVIDIA GPU:Broad support.
Fallback:PyTorch path.
PyPI:pip install.
Tags:llm

