llmfit

llmfit

Find which LLMs your hardware can run

Description

You want to run open models locally, but between 7B, 14B and endless quantizations you have no idea what your GPU and RAM can handle. llmfit detects your CPU, RAM, GPUs and VRAM and checks hundreds of models and quantizations to show what will run comfortably.

Its terminal UI lists the recommendations, can download a model and measure real tokens per second, and lets you contribute results so others on the same hardware get measured numbers. It supports NVIDIA, Apple Silicon, AMD and Intel.

Features



Hardware detection:CPU, RAM, discrete and integrated GPUs, VRAM and unified memory.

Model fit:Scores models by size, context length and quantization.

Real benchmarks:Download, serve and measure tok/s.

Shared results:Submit measurements for identical hardware.