Best LLM for 8 GB VRAM · 2026

Best LLM for 8 GB VRAM: what fits, and what to run

Updated 8 Oct 2026 1 new this month

A 8 GB card hands a runtime about 7.0 GB once the display and driver have taken their share. At Q4_K_M with an 8K context, 28 of 80 models in the catalogue fit — 22 with headroom, 6 tightly. The pick for most people is Gemma 4 E4B: Laptop class with images and audio in. Speeds are estimated for a GeForce RTX 4060; the fit verdicts are the same for every card in the class.

Which one should I run?

One pick per job, from the editorial shortlist, checked to fit at Q4_K_M and 8K. How picks are chosen.

RecommendationModelWhyQuant · totalTok/s est.
Best overall Gemma 4 E4B
Google · Apache 2.0
Laptop class with images and audio in. Q4_K_M · 5.2 GB 37 Runs great
Best for coding Qwen3.5 4B
Alibaba · Apache 2.0
A coding agent in 3.4 GB — the 6–8 GB answer. Q4_K_M · 3.5 GB 63 Runs great
Best reasoning Qwen3.5 4B
Alibaba · Apache 2.0
Step-by-step reasoning in 8 GB. Q4_K_M · 3.5 GB 63 Runs great
Best for writing Gemma 4 E4B
Google · Apache 2.0
Readable prose on a laptop. Q4_K_M · 5.2 GB 37 Runs great
Best vision Gemma 4 E4B
Google · Apache 2.0
Laptop vision, audio too. Q4_K_M · 5.2 GB 37 Runs great
Best for agents Fara 7B
Microsoft · MIT
Web computer-use only — clicks and fills forms from screenshots, and stops to ask before anything irreversible. Q4_K_M · 5.7 GB 35 Runs great
Best translation Gemma 4 E4B
Google · Apache 2.0
140+ languages on a laptop. Q4_K_M · 5.2 GB 37 Runs great
Fastest good model LFM2.5 8B-A1B
Liquid AI · LFM Open License v1.0
1.5B active; designed for laptops without a GPU. Q4_K_M · 5.5 GB 75 Runs great
Best long context Qwen3.5 2B
Alibaba · Apache 2.0
Long documents on 4 GB. Q4_K_M · 3.4 GB
at 128K context
129 Runs great

The long-context row is judged at 128K: Qwen3.5 2B holds that window on 8 GB at Q4_K_M with 1.50 GB of KV cache — that is why it can differ from the best overall pick.

Everything that fits 8 GB

Largest first. "Max context" is the longest window the card holds at that quantisation with an f16 cache; a q8_0 cache roughly doubles it.

ModelParamsQuantWeights+KV 8KTotalHeadroomMax contextTok/s est.
LFM2.5 8B-A1B 8.47B MoE Q4_K_M 4.8 GB 0.09 GB 5.5 GB 1.5 GB 125K 75 Runs great
Fara 7B 8.29B Q4_K_M 4.7 GB 0.44 GB 5.7 GB 1.3 GB 31K 35 Runs great
Gemma 4 E4B 8.0B Q4_K_M 4.5 GB 0.14 GB 5.2 GB 1.8 GB 128K 37 Runs great
Ling 3.0 Tiny 7.9B-A1.3B 7.9B MoE Q4_K_M 4.4 GB 0.05 GB 5.1 GB 1.9 GB 128K 87 Runs great
DeepSeek-R1-Distill-Qwen 7B 7.62B Q4_K_M 4.3 GB 0.44 GB 5.3 GB 1.7 GB 38K 38 Runs great
Qwen2.5-Coder 7B 7.62B Q4_K_M 4.3 GB 0.44 GB 5.3 GB 1.7 GB 38K 38 Runs great
Mistral 7B Instruct v0.3 7.25B Q4_K_M 4.1 GB 1.00 GB 5.7 GB 1.3 GB 18K 40 Runs great
Gemma 4 E2B 5.1B Q4_K_M 2.9 GB 0.07 GB 3.5 GB 3.5 GB 128K 57 Runs great
Qwen3.5 4B 4.66B Q4_K_M 2.6 GB 0.25 GB 3.5 GB 3.5 GB 120K 63 Runs great
Gemma 3 4B 4.3B Q4_K_M 2.4 GB 0.30 GB 3.3 GB 3.7 GB 128K 68 Runs great
Qwen3 4B 4.02B Q4_K_M 2.3 GB 1.13 GB 4.0 GB 3.0 GB 29K 73 Runs great
Ministral 3 3B 3.85B Q4_K_M 2.2 GB 0.81 GB 3.6 GB 3.4 GB 41K 76 Runs great
Phi-4-mini 3.8B 3.84B Q4_K_M 2.2 GB 1.00 GB 3.8 GB 3.2 GB 33K 76 Runs great
Granite 4.1 3B 3.4B Q4_K_M 1.9 GB 0.63 GB 3.1 GB 3.9 GB 57K 86 Runs great
Llama 3.2 3B Instruct 3.21B Q4_K_M 1.8 GB 0.88 GB 3.3 GB 3.7 GB 42K 91 Runs great
LFM2.5 2.6B 2.7B Q4_K_M 1.5 GB 0.13 GB 2.2 GB 4.8 GB 128K 108 Runs great
Qwen3.5 2B 2.27B Q4_K_M 1.3 GB 0.09 GB 2.0 GB 5.0 GB 256K 129 Runs great
Qwen3 1.7B 1.72B Q4_K_M 1.0 GB 0.88 GB 2.4 GB 4.6 GB 32K 170 Runs great
Llama 3.2 1B Instruct 1.24B Q4_K_M 0.7 GB 0.25 GB 1.5 GB 5.5 GB 128K 236 Runs great
Gemma 3 1B 1.0B Q4_K_M 0.6 GB 0.04 GB 1.2 GB 5.8 GB 32K 293 Runs great
Qwen3.5 0.8B 0.87B Q4_K_M 0.5 GB 0.09 GB 1.2 GB 5.8 GB 256K 337 Runs great
Qwen3 0.6B 0.6B Q4_K_M 0.3 GB 0.88 GB 1.8 GB 5.2 GB 32K 488 Runs great
Qwen3.5 9B 9.65B Q4_K_M 5.4 GB 0.25 GB 6.3 GB 0.7 GB 31K 30 Tight
Ornith 1.5 9B 9.41B Q4_K_M 5.3 GB 0.25 GB 6.1 GB 0.9 GB 35K 31 Tight
Ministral 3 8B 8.92B Q4_K_M 5.0 GB 1.06 GB 6.7 GB 0.3 GB 10K 33 Tight
Granite 4.1 8B 8.79B Q4_K_M 4.9 GB 1.25 GB 6.8 GB 0.2 GB 9K 33 Tight
Qwen3 8B 8.19B Q4_K_M 4.6 GB 1.13 GB 6.3 GB 0.7 GB 12K 36 Tight
Llama 3.1 8B Instruct 8.03B Q4_K_M 4.5 GB 1.00 GB 6.1 GB 0.9 GB 15K 36 Tight

Just out of reach

These run with part of the weights in system RAM (32 GB assumed), at a few tokens per second. Each needs a bigger card to run properly — the link says which.

ModelParamsTotal @ Q4_K_MOver budgetTok/s est., offloadedNeeds
Qwen3.6 35B-A3B 35.9B MoE 20.9 GB +13.9 GB ~9.9 32 GB card
Ornith 1.5 35B-A3B 35.9B MoE 20.9 GB +13.9 GB ~9.9 32 GB card
LLM-jp 4 33B Thinking 33.2B 21.3 GB +14.3 GB ~2.4 32 GB card
Qwen3 32B 32.8B 21.0 GB +14.0 GB ~2.4 32 GB card
DeepSeek-R1-Distill-Qwen 32B 32.8B 21.0 GB +14.0 GB ~2.4 32 GB card
Qwen2.5-Coder 32B 32.8B 21.0 GB +14.0 GB ~2.4 32 GB card

All VRAM classes 12 GB → Adjust context, quant or cache →

Every total is quantised weights + f16 KV cache at 8K + 0.6 GB runtime overhead, against 7.0 GB usable. Tokens per second are estimated from memory bandwidth and never measured. The method.