AirLLM

AirLLM

70B inference with a single 4GB GPU

Description

Running a 70B model locally with just a 4GB GPU seems impossible. AirLLM runs 70B LLM inference on a single 4GB GPU without quantization, distillation or pruning.

It loads layers one at a time to slash memory, supports Llama, Qwen and more, and runs on Mac too.

Features



Low VRAM:70B on 4GB.

Layer-wise:Sequential loading.

No quantization:Full precision.

Models:Llama and Qwen.