LLM VRAM calculator
LLM VRAM calculator: how much VRAM does a model need on your GPU?
Pick a model and a device and you get the whole memory bill — quantised weights, KV cache at your context length, runtime overhead — against what the device can actually use, with a verdict, the longest context that fits, an estimated speed and the command to run it. Every number is computed from the model's published shape; nothing is benchmarked.
How the number is built
VRAM needed = weights + KV cache + overhead, checked against usable memory, which is the advertised size less what the display and driver keep (about 1.6 GB on a discrete card; a quarter of a Mac's unified memory).
- Weights = parameters × bits per weight ÷ 8. At Q4_K_M (4.83 effective bits) that is about 0.6 GB per billion parameters; at Q8_0, 1.06 GB. For a mixture-of-experts model every parameter counts, not just the active ones — they all have to be resident.
- KV cache = layers × 2 × KV heads × head dimension × context × bytes per element. It is linear in context, so 128K costs sixteen times 8K — unless the model uses sliding-window, latent or hybrid attention, which the calculator models exactly from each config. Its own calculator.
- Overhead = a flat 0.6 GB for the CUDA or Metal context and compute buffers. The least precise number here.
A fit with at least 15% of usable memory spare is "runs great"; under that it is "tight"; past it, the overflow streams from system RAM at a few tokens per second. The method, in full.
Quick reference: how much VRAM popular models need
At the recommended quantisation with an f16 cache. "Total" includes 0.6 GB overhead; "needs" is the smallest discrete class with 15% headroom at 8K.
| Model | Quant | Weights | KV @ 8K | Total @ 8K | Total @ 32K | Needs |
|---|---|---|---|---|---|---|
| Qwen3.5 4B 4.66B |
Q4_K_M | 2.6 GB | 0.25 GB | 3.5 GB | 4.2 GB | 8 GB card |
| Qwen3.5 9B 9.65B |
Q4_K_M | 5.4 GB | 0.25 GB | 6.3 GB | 7.0 GB | 10 GB card |
| Gemma 4 12B 12B |
Q4_K_M | 6.7 GB | 0.81 GB | 8.2 GB | 9.7 GB | 11 GB card |
| gpt-oss 20B 20.9B MoE · 3.6B active |
MXFP4 | 10.8 GB | 0.19 GB | 11.6 GB | 12.2 GB | 16 GB card |
| Mistral Small 3.2 24B 23.6B |
Q4_K_M | 13.3 GB | 1.25 GB | 15.1 GB | 18.9 GB | 20 GB card |
| Qwen3.8 27B 27.8B |
Q4_K_M | 15.6 GB | 0.50 GB | 16.7 GB | 18.2 GB | 24 GB card |
| Gemma 4 31B 31.3B |
Q4_K_M | 17.6 GB | 2.03 GB | 20.2 GB | 24.0 GB | 32 GB card |
| Qwen3.6 35B-A3B 35.9B MoE · 3.3B active |
Q4_K_M | 20.2 GB | 0.16 GB | 20.9 GB | 21.4 GB | 32 GB card |
| Llama 3.3 70B Instruct 70.6B |
Q4_K_M | 39.7 GB | 2.50 GB | 42.8 GB | 50.3 GB | 64 GB card |
| gpt-oss 120B 117B MoE · 5.1B active |
MXFP4 | 60.5 GB | 0.29 GB | 61.4 GB | 62.2 GB | 80 GB card |
| GLM-5.3-Flash 320B-A18B 321B MoE · 18B active |
Q4_K_M | 180.5 GB | 0.09 GB | 181.2 GB | 181.4 GB | no single card |
| DeepSeek V4.1 Flash 552B-A16B 552B MoE · 16B active |
Q4_K_M | 310.4 GB | 0.04 GB | 311.0 GB | 311.1 GB | no single card |
Reverse it
Every model page has a what hardware do I need? companion that checks all 91 devices at every quantisation. For Qwen3.8 27B.
By class
Best local LLM for 8–128 GB turns the same arithmetic into a pick per job.
The letters
GGUF and quantization explained — what Q4_K_M means and which file to download.