Caveman Compression

Caveman Compression

Semantic compression for LLM contexts

Description

Long documents in context waste tokens on predictable grammar words. Caveman Compression is a semantic compression method for LLM contexts.

It strips predictable grammar while keeping the unpredictable factual content, so models still understand with far fewer tokens.

Features



Semantic:Facts kept.

Tokens:Grammar removed.

Readable:Meaning intact.