
Description
Letting reasoning models think at length blows up the KV cache and overwhelms small GPUs. TriAttention from MIT, NVIDIA and ZJU compresses the KV cache with a trigonometric method, 10.7x smaller with 2.5x throughput on long reasoning and no accuracy loss.
It lets apps like OpenClaw run on memory-constrained local GPUs.
KV compression:10.7x.
Throughput:2.5x.
Lossless:Same accuracy.
Small GPUs:Runs locally.
It lets apps like OpenClaw run on memory-constrained local GPUs.
Features
KV compression:10.7x.
Throughput:2.5x.
Lossless:Same accuracy.
Small GPUs:Runs locally.
