Local LLM VRAM estimator
Local LLM VRAM planner
Estimate a VRAM planning target for local inference. It answers capacity, not model speed or tokens per second.
The default 7B, 4-bit values (2 GiB KV cache, 1 GiB runtime, and 20% headroom) are illustrative starting points, not representative requirements for every 7B model.
Planning assumptions: ideal weight memory is parameters × bits per weight ÷ 8, converted to GiB. KV cache and runtime overhead are allowances you enter; they are not measured by this page. KV-cache needs vary with architecture, context length, batch size, and parallelism, so they cannot be inferred from parameter count alone.
This estimate excludes quantization metadata, mixed precision, temporary buffers, model-specific kernels, operating-system display use, and any guarantee that a model will fit.
Background: Hugging Face quantization documentation; Hugging Face cache explanation; llama.cpp performance notes.
Compare observed prices: 16GB+ GPUs · NVIDIA 24GB+ · 32GB+ GPUs. Check your runtime's actual requirements before choosing a card.