Local LLM VRAM estimator

Local LLM VRAM planner

Estimate a VRAM planning target for local inference. It answers capacity, not model speed or tokens per second.

Plan for 7.51 GiB total VRAM (3.26 GiB ideal weights + 2 GiB KV cache + 1 GiB runtime, then 20% headroom).

The default 7B, 4-bit values (2 GiB KV cache, 1 GiB runtime, and 20% headroom) are illustrative starting points, not representative requirements for every 7B model.

Planning assumptions: ideal weight memory is parameters × bits per weight ÷ 8, converted to GiB. KV cache and runtime overhead are allowances you enter; they are not measured by this page. KV-cache needs vary with architecture, context length, batch size, and parallelism, so they cannot be inferred from parameter count alone.

This estimate excludes quantization metadata, mixed precision, temporary buffers, model-specific kernels, operating-system display use, and any guarantee that a model will fit.

Background: Hugging Face quantization documentation; Hugging Face cache explanation; llama.cpp performance notes.

Compare observed prices: 16GB+ GPUs · NVIDIA 24GB+ · 32GB+ GPUs. Check your runtime's actual requirements before choosing a card.