Best LLM for 8 GB VRAM · 2026
Best LLM for 8 GB VRAM: what fits, and what to run
A 8 GB card hands a runtime about 7.0 GB once the display and driver have taken their share. At Q4_K_M with an 8K context, 28 of 80 models in the catalogue fit — 22 with headroom, 6 tightly. The pick for most people is Gemma 4 E4B: Laptop class with images and audio in. Speeds are estimated for a GeForce RTX 4060; the fit verdicts are the same for every card in the class.
Which one should I run?
One pick per job, from the editorial shortlist, checked to fit at Q4_K_M and 8K. How picks are chosen.
| Recommendation | Model | Why | Quant · total | Tok/s est. | |
|---|---|---|---|---|---|
| Best overall | Gemma 4 E4B Google · Apache 2.0 |
Laptop class with images and audio in. | Q4_K_M · 5.2 GB | 37 | Runs great |
| Best for coding | Qwen3.5 4B Alibaba · Apache 2.0 |
A coding agent in 3.4 GB — the 6–8 GB answer. | Q4_K_M · 3.5 GB | 63 | Runs great |
| Best reasoning | Qwen3.5 4B Alibaba · Apache 2.0 |
Step-by-step reasoning in 8 GB. | Q4_K_M · 3.5 GB | 63 | Runs great |
| Best for writing | Gemma 4 E4B Google · Apache 2.0 |
Readable prose on a laptop. | Q4_K_M · 5.2 GB | 37 | Runs great |
| Best vision | Gemma 4 E4B Google · Apache 2.0 |
Laptop vision, audio too. | Q4_K_M · 5.2 GB | 37 | Runs great |
| Best for agents | Fara 7B Microsoft · MIT |
Web computer-use only — clicks and fills forms from screenshots, and stops to ask before anything irreversible. | Q4_K_M · 5.7 GB | 35 | Runs great |
| Best translation | Gemma 4 E4B Google · Apache 2.0 |
140+ languages on a laptop. | Q4_K_M · 5.2 GB | 37 | Runs great |
| Fastest good model | LFM2.5 8B-A1B Liquid AI · LFM Open License v1.0 |
1.5B active; designed for laptops without a GPU. | Q4_K_M · 5.5 GB | 75 | Runs great |
| Best long context | Qwen3.5 2B Alibaba · Apache 2.0 |
Long documents on 4 GB. | Q4_K_M · 3.4 GB at 128K context |
129 | Runs great |
The long-context row is judged at 128K: Qwen3.5 2B holds that window on 8 GB at Q4_K_M with 1.50 GB of KV cache — that is why it can differ from the best overall pick.
Everything that fits 8 GB
Largest first. "Max context" is the longest window the card holds at that quantisation with an f16 cache; a q8_0 cache roughly doubles it.
| Model | Params | Quant | Weights | +KV 8K | Total | Headroom | Max context | Tok/s est. | |
|---|---|---|---|---|---|---|---|---|---|
| LFM2.5 8B-A1B | 8.47B MoE | Q4_K_M | 4.8 GB | 0.09 GB | 5.5 GB | 1.5 GB | 125K | 75 | Runs great |
| Fara 7B | 8.29B | Q4_K_M | 4.7 GB | 0.44 GB | 5.7 GB | 1.3 GB | 31K | 35 | Runs great |
| Gemma 4 E4B | 8.0B | Q4_K_M | 4.5 GB | 0.14 GB | 5.2 GB | 1.8 GB | 128K | 37 | Runs great |
| Ling 3.0 Tiny 7.9B-A1.3B | 7.9B MoE | Q4_K_M | 4.4 GB | 0.05 GB | 5.1 GB | 1.9 GB | 128K | 87 | Runs great |
| DeepSeek-R1-Distill-Qwen 7B | 7.62B | Q4_K_M | 4.3 GB | 0.44 GB | 5.3 GB | 1.7 GB | 38K | 38 | Runs great |
| Qwen2.5-Coder 7B | 7.62B | Q4_K_M | 4.3 GB | 0.44 GB | 5.3 GB | 1.7 GB | 38K | 38 | Runs great |
| Mistral 7B Instruct v0.3 | 7.25B | Q4_K_M | 4.1 GB | 1.00 GB | 5.7 GB | 1.3 GB | 18K | 40 | Runs great |
| Gemma 4 E2B | 5.1B | Q4_K_M | 2.9 GB | 0.07 GB | 3.5 GB | 3.5 GB | 128K | 57 | Runs great |
| Qwen3.5 4B | 4.66B | Q4_K_M | 2.6 GB | 0.25 GB | 3.5 GB | 3.5 GB | 120K | 63 | Runs great |
| Gemma 3 4B | 4.3B | Q4_K_M | 2.4 GB | 0.30 GB | 3.3 GB | 3.7 GB | 128K | 68 | Runs great |
| Qwen3 4B | 4.02B | Q4_K_M | 2.3 GB | 1.13 GB | 4.0 GB | 3.0 GB | 29K | 73 | Runs great |
| Ministral 3 3B | 3.85B | Q4_K_M | 2.2 GB | 0.81 GB | 3.6 GB | 3.4 GB | 41K | 76 | Runs great |
| Phi-4-mini 3.8B | 3.84B | Q4_K_M | 2.2 GB | 1.00 GB | 3.8 GB | 3.2 GB | 33K | 76 | Runs great |
| Granite 4.1 3B | 3.4B | Q4_K_M | 1.9 GB | 0.63 GB | 3.1 GB | 3.9 GB | 57K | 86 | Runs great |
| Llama 3.2 3B Instruct | 3.21B | Q4_K_M | 1.8 GB | 0.88 GB | 3.3 GB | 3.7 GB | 42K | 91 | Runs great |
| LFM2.5 2.6B | 2.7B | Q4_K_M | 1.5 GB | 0.13 GB | 2.2 GB | 4.8 GB | 128K | 108 | Runs great |
| Qwen3.5 2B | 2.27B | Q4_K_M | 1.3 GB | 0.09 GB | 2.0 GB | 5.0 GB | 256K | 129 | Runs great |
| Qwen3 1.7B | 1.72B | Q4_K_M | 1.0 GB | 0.88 GB | 2.4 GB | 4.6 GB | 32K | 170 | Runs great |
| Llama 3.2 1B Instruct | 1.24B | Q4_K_M | 0.7 GB | 0.25 GB | 1.5 GB | 5.5 GB | 128K | 236 | Runs great |
| Gemma 3 1B | 1.0B | Q4_K_M | 0.6 GB | 0.04 GB | 1.2 GB | 5.8 GB | 32K | 293 | Runs great |
| Qwen3.5 0.8B | 0.87B | Q4_K_M | 0.5 GB | 0.09 GB | 1.2 GB | 5.8 GB | 256K | 337 | Runs great |
| Qwen3 0.6B | 0.6B | Q4_K_M | 0.3 GB | 0.88 GB | 1.8 GB | 5.2 GB | 32K | 488 | Runs great |
| Qwen3.5 9B | 9.65B | Q4_K_M | 5.4 GB | 0.25 GB | 6.3 GB | 0.7 GB | 31K | 30 | Tight |
| Ornith 1.5 9B | 9.41B | Q4_K_M | 5.3 GB | 0.25 GB | 6.1 GB | 0.9 GB | 35K | 31 | Tight |
| Ministral 3 8B | 8.92B | Q4_K_M | 5.0 GB | 1.06 GB | 6.7 GB | 0.3 GB | 10K | 33 | Tight |
| Granite 4.1 8B | 8.79B | Q4_K_M | 4.9 GB | 1.25 GB | 6.8 GB | 0.2 GB | 9K | 33 | Tight |
| Qwen3 8B | 8.19B | Q4_K_M | 4.6 GB | 1.13 GB | 6.3 GB | 0.7 GB | 12K | 36 | Tight |
| Llama 3.1 8B Instruct | 8.03B | Q4_K_M | 4.5 GB | 1.00 GB | 6.1 GB | 0.9 GB | 15K | 36 | Tight |
Just out of reach
These run with part of the weights in system RAM (32 GB assumed), at a few tokens per second. Each needs a bigger card to run properly — the link says which.
| Model | Params | Total @ Q4_K_M | Over budget | Tok/s est., offloaded | Needs |
|---|---|---|---|---|---|
| Qwen3.6 35B-A3B | 35.9B MoE | 20.9 GB | +13.9 GB | ~9.9 | 32 GB card |
| Ornith 1.5 35B-A3B | 35.9B MoE | 20.9 GB | +13.9 GB | ~9.9 | 32 GB card |
| LLM-jp 4 33B Thinking | 33.2B | 21.3 GB | +14.3 GB | ~2.4 | 32 GB card |
| Qwen3 32B | 32.8B | 21.0 GB | +14.0 GB | ~2.4 | 32 GB card |
| DeepSeek-R1-Distill-Qwen 32B | 32.8B | 21.0 GB | +14.0 GB | ~2.4 | 32 GB card |
| Qwen2.5-Coder 32B | 32.8B | 21.0 GB | +14.0 GB | ~2.4 | 32 GB card |
Every total is quantised weights + f16 KV cache at 8K + 0.6 GB runtime overhead, against 7.0 GB usable. Tokens per second are estimated from memory bandwidth and never measured. The method.