
Description
Running a 70B model locally with just a 4GB GPU seems impossible. AirLLM runs 70B LLM inference on a single 4GB GPU without quantization, distillation or pruning.
It loads layers one at a time to slash memory, supports Llama, Qwen and more, and runs on Mac too.
Low VRAM:70B on 4GB.
Layer-wise:Sequential loading.
No quantization:Full precision.
Models:Llama and Qwen.
It loads layers one at a time to slash memory, supports Llama, Qwen and more, and runs on Mac too.
Features
Low VRAM:70B on 4GB.
Layer-wise:Sequential loading.
No quantization:Full precision.
Models:Llama and Qwen.
