LLM VRAM calculator

LLM VRAM calculator: how much VRAM does a model need on your GPU?

Pick a model and a device and you get the whole memory bill — quantised weights, KV cache at your context length, runtime overhead — against what the device can actually use, with a verdict, the longest context that fits, an estimated speed and the command to run it. Every number is computed from the model's published shape; nothing is benchmarked.

Opens the sizing page for that pair, with every quantisation laddered.

How the number is built

VRAM needed = weights + KV cache + overhead, checked against usable memory, which is the advertised size less what the display and driver keep (about 1.6 GB on a discrete card; a quarter of a Mac's unified memory).

  • Weights = parameters × bits per weight ÷ 8. At Q4_K_M (4.83 effective bits) that is about 0.6 GB per billion parameters; at Q8_0, 1.06 GB. For a mixture-of-experts model every parameter counts, not just the active ones — they all have to be resident.
  • KV cache = layers × 2 × KV heads × head dimension × context × bytes per element. It is linear in context, so 128K costs sixteen times 8K — unless the model uses sliding-window, latent or hybrid attention, which the calculator models exactly from each config. Its own calculator.
  • Overhead = a flat 0.6 GB for the CUDA or Metal context and compute buffers. The least precise number here.

A fit with at least 15% of usable memory spare is "runs great"; under that it is "tight"; past it, the overflow streams from system RAM at a few tokens per second. The method, in full.

Quick reference: how much VRAM popular models need

At the recommended quantisation with an f16 cache. "Total" includes 0.6 GB overhead; "needs" is the smallest discrete class with 15% headroom at 8K.

ModelQuantWeightsKV @ 8KTotal @ 8KTotal @ 32KNeeds
Qwen3.5 4B
4.66B
Q4_K_M 2.6 GB 0.25 GB 3.5 GB 4.2 GB 8 GB card
Qwen3.5 9B
9.65B
Q4_K_M 5.4 GB 0.25 GB 6.3 GB 7.0 GB 10 GB card
Gemma 4 12B
12B
Q4_K_M 6.7 GB 0.81 GB 8.2 GB 9.7 GB 11 GB card
gpt-oss 20B
20.9B MoE · 3.6B active
MXFP4 10.8 GB 0.19 GB 11.6 GB 12.2 GB 16 GB card
Mistral Small 3.2 24B
23.6B
Q4_K_M 13.3 GB 1.25 GB 15.1 GB 18.9 GB 20 GB card
Qwen3.8 27B
27.8B
Q4_K_M 15.6 GB 0.50 GB 16.7 GB 18.2 GB 24 GB card
Gemma 4 31B
31.3B
Q4_K_M 17.6 GB 2.03 GB 20.2 GB 24.0 GB 32 GB card
Qwen3.6 35B-A3B
35.9B MoE · 3.3B active
Q4_K_M 20.2 GB 0.16 GB 20.9 GB 21.4 GB 32 GB card
Llama 3.3 70B Instruct
70.6B
Q4_K_M 39.7 GB 2.50 GB 42.8 GB 50.3 GB 64 GB card
gpt-oss 120B
117B MoE · 5.1B active
MXFP4 60.5 GB 0.29 GB 61.4 GB 62.2 GB 80 GB card
GLM-5.3-Flash 320B-A18B
321B MoE · 18B active
Q4_K_M 180.5 GB 0.09 GB 181.2 GB 181.4 GB no single card
DeepSeek V4.1 Flash 552B-A16B
552B MoE · 16B active
Q4_K_M 310.4 GB 0.04 GB 311.0 GB 311.1 GB no single card

Reverse it

Every model page has a what hardware do I need? companion that checks all 91 devices at every quantisation. For Qwen3.8 27B.

By class

Best local LLM for 8–128 GB turns the same arithmetic into a pick per job.

The letters

GGUF and quantization explained — what Q4_K_M means and which file to download.