KV cache calculator

KV cache calculator: how much memory does context cost?

The KV cache is the part of an LLM's memory bill that grows with every token you keep in the window — and the reason a model that "should fit" does not. This page computes it per token and per context length for every model in the catalogue, from each one's published layer count, attention shape and cache type, including sliding-window, latent and hybrid attention.

bytes per token = caching layers × 2 × KV heads × head dim × bytes per element — then multiply by the context length. At f16 that is 2 bytes per element; a q8_0 cache is 1.0625 and roughly halves it for a negligible loss; q4_0 quarters it and degrades long-context recall. Three families cache far less than the formula suggests and the table says which: sliding-window models keep only a few hundred tokens in most layers, latent-attention models (DeepSeek, GLM, Kimi) cache a compressed vector, and the 2026 hybrids (Qwen3.5+, Nemotron 3, LFM2.5) cache in only a fraction of their blocks.

Phi-4-mini 3.8B: 128.0 KB per token at f16

Phi-4-mini 3.8B has 32 blocks, 8 KV heads and a head dimension of 128. Per token, per caching layer: 2 × 8 × 128 = 2048 elements × 2 bytes.

ContextKV cache (f16)Total with Q4_K_M weightsNeeds
4K0.50 GB3.3 GB8 GB card
8K1.00 GB3.8 GB8 GB card
32K4.00 GB6.8 GB10 GB card
128K (max)16.00 GB18.8 GB24 GB card

Total adds 2.2 GB of weights and 0.6 GB overhead. What hardware Phi-4-mini 3.8B needs · check it on a device.

Every model, per token and per context

At f16. Per-token figures are for a full-attention layer stack; the context columns apply each model's sliding-window saving. Click a model for its full table.

ModelParamsPer token@ 8K@ 32K@ 128KWeights @ Q4
Qwen3 0.6B
full GQA
0.6B 112.0 KB 0.88 GB 3.50 GB max 32K 0.3 GB
Qwen3.5 0.8B
hybrid: 6 of 24 blocks
0.87B 12.0 KB 0.09 GB 0.38 GB 1.5 GB 0.5 GB
Gemma 3 1B
sliding window 512
1.0B 15.0 KB 0.04 GB 0.14 GB max 32K 0.6 GB
Llama 3.2 1B Instruct
full GQA
1.24B 32.0 KB 0.25 GB 1.00 GB 4.0 GB 0.7 GB
Qwen3 1.7B
full GQA
1.72B 112.0 KB 0.88 GB 3.50 GB max 32K 1.0 GB
Qwen3.5 2B
hybrid: 6 of 24 blocks
2.27B 12.0 KB 0.09 GB 0.38 GB 1.5 GB 1.3 GB
LFM2.5 2.6B
hybrid: 8 of 30 blocks
2.7B 16.0 KB 0.13 GB 0.50 GB 2.0 GB 1.5 GB
Llama 3.2 3B Instruct
full GQA
3.21B 112.0 KB 0.88 GB 3.50 GB 14.0 GB 1.8 GB
Granite 4.1 3B
full GQA
3.4B 80.0 KB 0.63 GB 2.50 GB 10.0 GB 1.9 GB
Phi-4-mini 3.8B
full GQA
3.84B 128.0 KB 1.00 GB 4.00 GB 16.0 GB 2.2 GB
Ministral 3 3B
full GQA
3.85B 104.0 KB 0.81 GB 3.25 GB 13.0 GB 2.2 GB
Qwen3 4B
full GQA
4.02B 144.0 KB 1.13 GB 4.50 GB max 32K 2.3 GB
Gemma 3 4B
sliding window 1024
4.3B 136.0 KB 0.30 GB 0.86 GB 3.1 GB 2.4 GB
Qwen3.5 4B
hybrid: 8 of 32 blocks
4.66B 32.0 KB 0.25 GB 1.00 GB 4.0 GB 2.6 GB
Gemma 4 E2B
sliding window 512
5.1B 21.0 KB 0.07 GB 0.23 GB 0.9 GB 2.9 GB
Mistral 7B Instruct v0.3
full GQA
7.25B 128.0 KB 1.00 GB 4.00 GB max 32K 4.1 GB
Olmo 3 7B Instruct
sliding window 4096
7.3B 512.0 KB 2.50 GB 5.50 GB max 64K 4.1 GB
DeepSeek-R1-Distill-Qwen 7B
full GQA
7.62B 56.0 KB 0.44 GB 1.75 GB 7.0 GB 4.3 GB
Qwen2.5-Coder 7B
full GQA
7.62B 56.0 KB 0.44 GB 1.75 GB 7.0 GB 4.3 GB
Ling 3.0 Tiny 7.9B-A1.3B
latent (MLA)
7.9B MoE 6.8 KB 0.05 GB 0.21 GB 0.8 GB 4.4 GB
Gemma 4 E4B
sliding window 512
8.0B 49.0 KB 0.14 GB 0.47 GB 1.8 GB 4.5 GB
Llama 3.1 8B Instruct
full GQA
8.03B 128.0 KB 1.00 GB 4.00 GB 16.0 GB 4.5 GB
Qwen3 8B
full GQA
8.19B 144.0 KB 1.13 GB 4.50 GB 18.0 GB 4.6 GB
Fara 7B
full GQA
8.29B 56.0 KB 0.44 GB 1.75 GB max 125K 4.7 GB
LFM2.5 8B-A1B
hybrid: 6 of 24 blocks
8.47B MoE 12.0 KB 0.09 GB 0.38 GB max 125K 4.8 GB
Granite 4.1 8B
full GQA
8.79B 160.0 KB 1.25 GB 5.00 GB 20.0 GB 4.9 GB
Ministral 3 8B
full GQA
8.92B 136.0 KB 1.06 GB 4.25 GB 17.0 GB 5.0 GB
Ornith 1.5 9B
hybrid: 8 of 32 blocks
9.41B 32.0 KB 0.25 GB 1.00 GB 4.0 GB 5.3 GB
Qwen3.5 9B
hybrid: 8 of 32 blocks
9.65B 32.0 KB 0.25 GB 1.00 GB 4.0 GB 5.4 GB
Gemma 4 12B
sliding window 1024
12B 384.0 KB 0.81 GB 2.31 GB 8.3 GB 6.7 GB
Gemma 3 12B
sliding window 1024
12.2B 384.0 KB 0.81 GB 2.31 GB 8.3 GB 6.9 GB
Mistral NeMo 12B
full GQA
12.2B 160.0 KB 1.25 GB 5.00 GB 20.0 GB 6.9 GB
Ministral 3 14B
full GQA
13.9B 160.0 KB 1.25 GB 5.00 GB 20.0 GB 7.8 GB
Phi-4 14B
full GQA
14.7B 200.0 KB 1.56 GB 3.13 GB max 16K 8.3 GB
Qwen3 14B
full GQA
14.8B 160.0 KB 1.25 GB 5.00 GB 20.0 GB 8.3 GB
DeepSeek-R1-Distill-Qwen 14B
full GQA
14.8B 192.0 KB 1.50 GB 6.00 GB 24.0 GB 8.3 GB
Qwen2.5-Coder 14B
full GQA
14.8B 192.0 KB 1.50 GB 6.00 GB 24.0 GB 8.3 GB
gpt-oss 20B
sliding window 128
20.9B MoE 27.0 KB 0.19 GB 0.75 GB 3.0 GB 10.8 GB
Mistral Small 3.2 24B
full GQA
23.6B 160.0 KB 1.25 GB 5.00 GB 20.0 GB 13.3 GB
Devstral Small 2 24B
full GQA
24B 160.0 KB 1.25 GB 5.00 GB 20.0 GB 13.5 GB
Gemma 4 26B-A4B
sliding window 1024
26.5B MoE 240.0 KB 0.51 GB 1.45 GB 5.2 GB 14.9 GB
Gemma 3 27B
sliding window 1024
27.4B 496.0 KB 1.03 GB 2.91 GB 10.4 GB 15.4 GB
Qwen3.8 27B
hybrid: 16 of 64 blocks
27.8B 64.0 KB 0.50 GB 2.00 GB 8.0 GB 15.6 GB
Qwen3.6 27B
hybrid: 16 of 64 blocks
27.8B 64.0 KB 0.50 GB 2.00 GB 8.0 GB 15.6 GB
Granite 4.1 30B
full GQA
28.9B 256.0 KB 2.00 GB 8.00 GB 32.0 GB 16.3 GB
Muse Glimmer 30B
sliding window 2048
29.8B 52.0 KB 0.18 GB 0.48 GB 1.7 GB 16.8 GB
Qwen3 30B-A3B
full GQA
30.5B MoE 96.0 KB 0.75 GB 3.00 GB 12.0 GB 17.1 GB
Qwen3 Coder 30B-A3B
full GQA
30.5B MoE 96.0 KB 0.75 GB 3.00 GB 12.0 GB 17.1 GB
GLM-4.7-Flash 30B-A3B
latent (MLA)
31.2B MoE 52.9 KB 0.41 GB 1.65 GB 6.6 GB 17.5 GB
Gemma 4 31B
sliding window 1024
31.3B 960.0 KB 2.03 GB 5.78 GB 20.8 GB 17.6 GB
Nemotron 3.5 Lightning 30B-A3B
hybrid: 6 of 52 blocks
31.6B MoE 6.0 KB 0.05 GB 0.19 GB 0.8 GB 17.8 GB
Olmo 3.1 32B Instruct
sliding window 4096
32.2B 256.0 KB 1.25 GB 2.75 GB max 64K 18.1 GB
Qwen3 32B
full GQA
32.8B 256.0 KB 2.00 GB 8.00 GB 32.0 GB 18.4 GB
DeepSeek-R1-Distill-Qwen 32B
full GQA
32.8B 256.0 KB 2.00 GB 8.00 GB 32.0 GB 18.4 GB
Qwen2.5-Coder 32B
full GQA
32.8B 256.0 KB 2.00 GB 8.00 GB 32.0 GB 18.4 GB
LLM-jp 4 33B Thinking
full GQA
33.2B 256.0 KB 2.00 GB 8.00 GB max 64K 18.7 GB
Qwen3.6 35B-A3B
hybrid: 10 of 40 blocks
35.9B MoE 20.0 KB 0.16 GB 0.63 GB 2.5 GB 20.2 GB
Ornith 1.5 35B-A3B
hybrid: 10 of 40 blocks
35.9B MoE 20.0 KB 0.16 GB 0.63 GB 2.5 GB 20.2 GB
Llama 3.3 70B Instruct
full GQA
70.6B 320.0 KB 2.50 GB 10.00 GB 40.0 GB 39.7 GB
DeepSeek-R1-Distill-Llama 70B
full GQA
70.6B 320.0 KB 2.50 GB 10.00 GB 40.0 GB 39.7 GB
Llama 4 Scout 109B-A17B
sliding window 8192
109B MoE 192.0 KB 1.50 GB 2.63 GB 7.1 GB 61.3 GB
gpt-oss 120B
sliding window 128
117B MoE 40.5 KB 0.29 GB 1.13 GB 4.5 GB 60.5 GB
Mistral Small 4 119B-A6B
latent (MLA)
119B MoE 22.5 KB 0.18 GB 0.70 GB 2.8 GB 66.9 GB
Nemotron 3 Super 120B-A12B
hybrid: 8 of 88 blocks
124B MoE 8.0 KB 0.06 GB 0.25 GB 1.0 GB 69.7 GB
Ling 3.0 Flash 124B-A5B
latent (MLA)
124B MoE 7.9 KB 0.06 GB 0.25 GB 1.0 GB 69.7 GB
Qwen3.5 122B-A10B
hybrid: 12 of 48 blocks
125B MoE 24.0 KB 0.19 GB 0.75 GB 3.0 GB 70.3 GB
Qwen3.8-Flash-Next 180B-A6B
hybrid: 12 of 48 blocks
180B MoE 24.0 KB 0.19 GB 0.75 GB 3.0 GB 101.2 GB
Qwen3 235B-A22B
full GQA
235B MoE 188.0 KB 1.47 GB 5.88 GB 23.5 GB 132.1 GB
DeepSeek V4 Flash 284B-A13B
latent (MLA)
284B MoE 48.4 KB 0.38 GB 1.51 GB 6.0 GB 159.7 GB
GLM-5.3-Flash 320B-A18B
latent (MLA)
321B MoE 11.0 KB 0.09 GB 0.34 GB 1.4 GB 180.5 GB
Ornith 1.5 397B-A17B
hybrid: 15 of 60 blocks
397B MoE 30.0 KB 0.23 GB 0.94 GB 3.8 GB 223.2 GB
Qwen3.5 397B-A17B
hybrid: 15 of 60 blocks
403B MoE 30.0 KB 0.23 GB 0.94 GB 3.8 GB 226.6 GB
DeepSeek V4.1 Flash 552B-A16B
latent (MLA)
552B MoE 9.6 KB 0.04 GB 0.15 GB 0.6 GB 310.4 GB
DeepSeek-R1 671B
latent (MLA)
671B MoE 68.6 KB 0.54 GB 2.14 GB 8.6 GB 377.3 GB
GLM-5.2 744B-A40B
latent (MLA)
753B MoE 87.8 KB 0.69 GB 2.74 GB 11.0 GB 423.4 GB
GLM-5.3 744B-A40B
latent (MLA)
753B MoE 87.8 KB 0.69 GB 2.74 GB 11.0 GB 423.4 GB
Kimi K2.6 1T-A32B
latent (MLA)
1027B MoE 68.6 KB 0.54 GB 2.14 GB 8.6 GB 577.5 GB
DeepSeek V4 Pro 1.6T-A49B
latent (MLA)
1650B MoE 68.6 KB 0.54 GB 2.14 GB 8.6 GB 927.8 GB
Qwen3.8 2.4T-A95B
hybrid: 23 of 92 blocks
2446B MoE 92.0 KB 0.72 GB 2.88 GB 11.5 GB 1375.4 GB
Kimi K3 2.8T-A104B
latent (MLA)
2780B MoE 27.0 KB 0.21 GB 0.84 GB 3.4 GB 1563.2 GB

The other half

The VRAM calculator adds the weights and overhead and checks the total against your device.

Serving several users

Each open conversation holds its own cache: multiply the 8K figure by concurrent users. Self-hosting notes.

Why it matters

The method explains the KV arithmetic, with the three attention families that escape it.