TriAttention

TriAttention

Efficient long reasoning with trigonometric KV compression

Description

Letting reasoning models think at length blows up the KV cache and overwhelms small GPUs. TriAttention from MIT, NVIDIA and ZJU compresses the KV cache with a trigonometric method, 10.7x smaller with 2.5x throughput on long reasoning and no accuracy loss.

It lets apps like OpenClaw run on memory-constrained local GPUs.

Features



KV compression:10.7x.

Throughput:2.5x.

Lossless:Same accuracy.

Small GPUs:Runs locally.