Best local LLM · 2026
The best local LLM for your VRAM, from 8 GB to 128 GB
There is no single best local LLM — there is a best one for the memory you have. This page answers that for each VRAM class: one pick per job (general assistant, coding, reasoning, writing, vision, agents, translation, speed, long context), chosen from a hand-ordered shortlist and checked, by arithmetic, to fit at Q4_K_M with an 8K context. The ordering is editorial and says so; the sizes and speeds are computed, never benchmarked.
At a glance
| Memory | Best overall | Coding | Reasoning | Vision | Fastest | Models that fit |
|---|---|---|---|---|---|---|
| 8 GB VRAM GeForce RTX 4060 |
Gemma 4 E4B | Qwen3.5 4B | Qwen3.5 4B | Gemma 4 E4B | LFM2.5 8B-A1B | 28 |
| 12 GB VRAM GeForce RTX 3060 12 GB |
Gemma 4 12B | Qwen3.5 9B | Qwen3.5 9B | Gemma 4 12B | LFM2.5 8B-A1B | 37 |
| 16 GB VRAM GeForce RTX 5080 |
Gemma 4 12B | Qwen3.5 9B | gpt-oss 20B | Gemma 4 12B | gpt-oss 20B | 38 |
| 24 GB VRAM GeForce RTX 4090 |
Qwen3.8 27B | Qwen3.8 27B | Qwen3.8 27B | Qwen3.8 27B | Gemma 4 26B-A4B | 58 |
| 32 GB VRAM GeForce RTX 5090 |
Qwen3.8 27B | Qwen3.8 27B | Gemma 4 31B | Qwen3.8 27B | Qwen3.6 35B-A3B | 58 |
| 48 GB VRAM RTX 6000 Ada |
Qwen3.8 27B | Qwen3.8 27B | Gemma 4 31B | Qwen3.8 27B | Qwen3.6 35B-A3B | 60 |
| 64 GB unified memory M4 Max · 64 GB |
Qwen3.8 27B | Qwen3.8 27B | Gemma 4 31B | Qwen3.8 27B | Qwen3.6 35B-A3B | 60 |
| 128 GB unified memory M4 Max · 128 GB |
Qwen3.8 27B | Ling 3.0 Flash 124B-A5B | Ling 3.0 Flash 124B-A5B | Qwen3.8 27B | Qwen3.6 35B-A3B | 66 |
Picks at Q4_K_M, 8K context, f16 KV cache. Not sure which class you are in? Detect your machine, or find it in the catalogue — every device page runs the same picks against your exact card.
Best local LLM for 8 GB VRAM
7.0 GB usable · 28 of 80 models fitEntry-level and laptop cards. Small dense models and the first MoEs. Speeds below are estimated for a GeForce RTX 4060; the fit verdicts hold for every 8 GB card.
| Recommendation | Model | Why | Quant · total | Tok/s est. | |
|---|---|---|---|---|---|
| Best overall | Gemma 4 E4B Google · Apache 2.0 |
Laptop class with images and audio in. | Q4_K_M · 5.2 GB | 37 | Runs great |
| Best for coding | Qwen3.5 4B Alibaba · Apache 2.0 |
A coding agent in 3.4 GB — the 6–8 GB answer. | Q4_K_M · 3.5 GB | 63 | Runs great |
| Best reasoning | Qwen3.5 4B Alibaba · Apache 2.0 |
Step-by-step reasoning in 8 GB. | Q4_K_M · 3.5 GB | 63 | Runs great |
| Best for writing | Gemma 4 E4B Google · Apache 2.0 |
Readable prose on a laptop. | Q4_K_M · 5.2 GB | 37 | Runs great |
| Best vision | Gemma 4 E4B Google · Apache 2.0 |
Laptop vision, audio too. | Q4_K_M · 5.2 GB | 37 | Runs great |
| Best for agents | Fara 7B Microsoft · MIT |
Web computer-use only — clicks and fills forms from screenshots, and stops to ask before anything irreversible. | Q4_K_M · 5.7 GB | 35 | Runs great |
| Best translation | Gemma 4 E4B Google · Apache 2.0 |
140+ languages on a laptop. | Q4_K_M · 5.2 GB | 37 | Runs great |
| Fastest good model | LFM2.5 8B-A1B Liquid AI · LFM Open License v1.0 |
1.5B active; designed for laptops without a GPU. | Q4_K_M · 5.5 GB | 75 | Runs great |
| Best long context | Qwen3.5 2B Alibaba · Apache 2.0 |
Long documents on 4 GB. | Q4_K_M · 3.4 GB at 128K context |
129 | Runs great |
Best local LLM for 12 GB VRAM
10.6 GB usable · 37 of 80 models fitThe most common upgrade step, and the first class where a 12B runs with headroom. Speeds below are estimated for a GeForce RTX 3060 12 GB; the fit verdicts hold for every 12 GB card.
| Recommendation | Model | Why | Quant · total | Tok/s est. | |
|---|---|---|---|---|---|
| Best overall | Gemma 4 12B Google · Apache 2.0 |
The 12 GB generalist: text, image and audio in, 140+ languages. | Q4_K_M · 8.2 GB | 32 | Runs great |
| Best for coding | Qwen3.5 9B Alibaba · Apache 2.0 |
The best coding model for 8–12 GB cards. | Q4_K_M · 6.3 GB | 40 | Runs great |
| Best reasoning | Qwen3.5 9B Alibaba · Apache 2.0 |
Thinking mode on a 12 GB card — ahead of the 2025 R1 distills. | Q4_K_M · 6.3 GB | 40 | Runs great |
| Best for writing | Gemma 4 12B Google · Apache 2.0 |
The 12 GB writing pick; 140+ languages. | Q4_K_M · 8.2 GB | 32 | Runs great |
| Best vision | Gemma 4 12B Google · Apache 2.0 |
Image and audio understanding on 12 GB. | Q4_K_M · 8.2 GB | 32 | Runs great |
| Best for agents | Granite 4.1 8B IBM · Apache 2.0 |
Enterprise-grade tool calling in a dense 8B, no thinking overhead. | Q4_K_M · 6.8 GB | 44 | Runs great |
| Best translation | Gemma 4 12B Google · Apache 2.0 |
Broad language coverage on a 12 GB card. | Q4_K_M · 8.2 GB | 32 | Runs great |
| Fastest good model | LFM2.5 8B-A1B Liquid AI · LFM Open License v1.0 |
1.5B active; designed for laptops without a GPU. | Q4_K_M · 5.5 GB | 99 | Runs great |
| Best long context | Qwen3.5 4B Alibaba · Apache 2.0 |
262K on 8 GB. | Q4_K_M · 7.2 GB at 128K context |
83 | Runs great |
Best local LLM for 16 GB VRAM
14.4 GB usable · 38 of 80 models fitWhere the 20B-class mixture-of-experts models start to fit whole. Speeds below are estimated for a GeForce RTX 5080; the fit verdicts hold for every 16 GB card.
| Recommendation | Model | Why | Quant · total | Tok/s est. | |
|---|---|---|---|---|---|
| Best overall | Gemma 4 12B Google · Apache 2.0 |
The 12 GB generalist: text, image and audio in, 140+ languages. | Q4_K_M · 8.2 GB | 86 | Runs great |
| Best for coding | Qwen3.5 9B Alibaba · Apache 2.0 |
The best coding model for 8–12 GB cards. | Q4_K_M · 6.3 GB | 107 | Runs great |
| Best reasoning | gpt-oss 20B OpenAI · Apache 2.0 |
Low/medium/high reasoning effort in 13 GB. | MXFP4 · 11.6 GB | 120 | Runs great |
| Best for writing | Gemma 4 12B Google · Apache 2.0 |
The 12 GB writing pick; 140+ languages. | Q4_K_M · 8.2 GB | 86 | Runs great |
| Best vision | Gemma 4 12B Google · Apache 2.0 |
Image and audio understanding on 12 GB. | Q4_K_M · 8.2 GB | 86 | Runs great |
| Best for agents | gpt-oss 20B OpenAI · Apache 2.0 |
Harmony-format tool calling, native 4-bit. | MXFP4 · 11.6 GB | 120 | Runs great |
| Best translation | Gemma 4 12B Google · Apache 2.0 |
Broad language coverage on a 12 GB card. | Q4_K_M · 8.2 GB | 86 | Runs great |
| Fastest good model | gpt-oss 20B OpenAI · Apache 2.0 |
3.6B active, native 4-bit. | MXFP4 · 11.6 GB | 120 | Runs great |
| Best long context | Qwen3.5 9B Alibaba · Apache 2.0 |
262K native in 8–12 GB. | Q4_K_M · 10.0 GB at 128K context |
107 | Runs great |
Best local LLM for 24 GB VRAM
22.4 GB usable · 58 of 80 models fitThe sweet spot: current-generation dense 27B models with room for context. Speeds below are estimated for a GeForce RTX 4090; the fit verdicts hold for every 24 GB card.
| Recommendation | Model | Why | Quant · total | Tok/s est. | |
|---|---|---|---|---|---|
| Best overall | Qwen3.8 27B Alibaba · Apache 2.0 |
Best capability per gigabyte: a current-generation dense 27B with vision and 262K context. | Q4_K_M · 16.7 GB | 39 | Runs great |
| Best for coding | Qwen3.8 27B Alibaba · Apache 2.0 |
Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. | Q4_K_M · 16.7 GB | 39 | Runs great |
| Best reasoning | Qwen3.8 27B Alibaba · Apache 2.0 |
Thinking mode plus the best agentic reasoning at 24 GB. | Q4_K_M · 16.7 GB | 39 | Runs great |
| Best for writing | Gemma 4 26B-A4B Google · Apache 2.0 |
Gemma 4 prose at MoE speed, with room for a long draft in context. | Q4_K_M · 16.0 GB | 110 | Runs great |
| Best vision | Qwen3.8 27B Alibaba · Apache 2.0 |
Image and video in, OSWorld 84 — the strongest local vision-language model at 24 GB. | Q4_K_M · 16.7 GB | 39 | Runs great |
| Best for agents | Qwen3.8 27B Alibaba · Apache 2.0 |
Computer use and tool calling are what 3.8 was trained for (OSWorld 84). | Q4_K_M · 16.7 GB | 39 | Runs great |
| Best translation | Qwen3.8 27B Alibaba · Apache 2.0 |
119 languages, and Qwen translations read as idiomatic rather than literal. | Q4_K_M · 16.7 GB | 39 | Runs great |
| Fastest good model | Gemma 4 26B-A4B Google · Apache 2.0 |
3.8B active and a small file — the quickest Gemma 4 with vision. | Q4_K_M · 16.0 GB | 110 | Runs great |
| Best long context | Nemotron 3.5 Lightning 30B-A3B NVIDIA · OpenMDW-1.1 |
262K, and a Mamba hybrid whose cache barely grows. | Q4_K_M · 19.1 GB at 128K context |
130 | Tight fit |
Best local LLM for 32 GB VRAM
30.4 GB usable · 58 of 80 models fitDense 30Bs at Q6, or a 27B with a very long context. Speeds below are estimated for a GeForce RTX 5090; the fit verdicts hold for every 32 GB card.
| Recommendation | Model | Why | Quant · total | Tok/s est. | |
|---|---|---|---|---|---|
| Best overall | Qwen3.8 27B Alibaba · Apache 2.0 |
Best capability per gigabyte: a current-generation dense 27B with vision and 262K context. | Q4_K_M · 16.7 GB | 69 | Runs great |
| Best for coding | Qwen3.8 27B Alibaba · Apache 2.0 |
Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. | Q4_K_M · 16.7 GB | 69 | Runs great |
| Best reasoning | Gemma 4 31B Google · Apache 2.0 |
89% AIME — the strongest maths of any model under 32B. | Q4_K_M · 20.2 GB | 62 | Runs great |
| Best for writing | Gemma 4 31B Google · Apache 2.0 |
The cleanest, least verbose prose of the 2026 24–32 GB class. | Q4_K_M · 20.2 GB | 62 | Runs great |
| Best vision | Qwen3.8 27B Alibaba · Apache 2.0 |
Image and video in, OSWorld 84 — the strongest local vision-language model at 24 GB. | Q4_K_M · 16.7 GB | 69 | Runs great |
| Best for agents | Qwen3.8 27B Alibaba · Apache 2.0 |
Computer use and tool calling are what 3.8 was trained for (OSWorld 84). | Q4_K_M · 16.7 GB | 69 | Runs great |
| Best translation | Qwen3.8 27B Alibaba · Apache 2.0 |
119 languages, and Qwen translations read as idiomatic rather than literal. | Q4_K_M · 16.7 GB | 69 | Runs great |
| Fastest good model | Qwen3.6 35B-A3B Alibaba · Apache 2.0 |
3.3B active per token: the fastest model that still competes with dense 27Bs. | Q4_K_M · 20.9 GB | 225 | Runs great |
| Best long context | Qwen3.8 27B Alibaba · Apache 2.0 |
262K native, and only 16 of 64 blocks keep a KV cache — long context is cheap here. | Q4_K_M · 24.2 GB at 128K context |
69 | Runs great |
Best local LLM for 48 GB VRAM
46.4 GB usable · 60 of 80 models fitWorkstation class: 70B at Q4, or a 30B near-lossless. Speeds below are estimated for a RTX 6000 Ada; the fit verdicts hold for every 48 GB card.
| Recommendation | Model | Why | Quant · total | Tok/s est. | |
|---|---|---|---|---|---|
| Best overall | Qwen3.8 27B Alibaba · Apache 2.0 |
Best capability per gigabyte: a current-generation dense 27B with vision and 262K context. | Q4_K_M · 16.7 GB | 37 | Runs great |
| Best for coding | Qwen3.8 27B Alibaba · Apache 2.0 |
Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. | Q4_K_M · 16.7 GB | 37 | Runs great |
| Best reasoning | Gemma 4 31B Google · Apache 2.0 |
89% AIME — the strongest maths of any model under 32B. | Q4_K_M · 20.2 GB | 33 | Runs great |
| Best for writing | Gemma 4 31B Google · Apache 2.0 |
The cleanest, least verbose prose of the 2026 24–32 GB class. | Q4_K_M · 20.2 GB | 33 | Runs great |
| Best vision | Qwen3.8 27B Alibaba · Apache 2.0 |
Image and video in, OSWorld 84 — the strongest local vision-language model at 24 GB. | Q4_K_M · 16.7 GB | 37 | Runs great |
| Best for agents | Qwen3.8 27B Alibaba · Apache 2.0 |
Computer use and tool calling are what 3.8 was trained for (OSWorld 84). | Q4_K_M · 16.7 GB | 37 | Runs great |
| Best translation | Qwen3.8 27B Alibaba · Apache 2.0 |
119 languages, and Qwen translations read as idiomatic rather than literal. | Q4_K_M · 16.7 GB | 37 | Runs great |
| Fastest good model | Qwen3.6 35B-A3B Alibaba · Apache 2.0 |
3.3B active per token: the fastest model that still competes with dense 27Bs. | Q4_K_M · 20.9 GB | 120 | Runs great |
| Best long context | Qwen3.8 27B Alibaba · Apache 2.0 |
262K native, and only 16 of 64 blocks keep a KV cache — long context is cheap here. | Q4_K_M · 24.2 GB at 128K context |
37 | Runs great |
Best local LLM for 64 GB unified memory
48.0 GB usable · 60 of 80 models fitThe 64 GB Mac. Unified memory, so the GPU can address about three quarters of it. Speeds below are estimated for a M4 Max · 64 GB; the fit verdicts hold for every 64 GB machine.
| Recommendation | Model | Why | Quant · total | Tok/s est. | |
|---|---|---|---|---|---|
| Best overall | Qwen3.8 27B Alibaba · Apache 2.0 |
Best capability per gigabyte: a current-generation dense 27B with vision and 262K context. | Q4_K_M · 16.7 GB | 15 | Runs great |
| Best for coding | Qwen3.8 27B Alibaba · Apache 2.0 |
Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. | Q4_K_M · 16.7 GB | 15 | Runs great |
| Best reasoning | Gemma 4 31B Google · Apache 2.0 |
89% AIME — the strongest maths of any model under 32B. | Q4_K_M · 20.2 GB | 13 | Runs great |
| Best for writing | Gemma 4 31B Google · Apache 2.0 |
The cleanest, least verbose prose of the 2026 24–32 GB class. | Q4_K_M · 20.2 GB | 13 | Runs great |
| Best vision | Qwen3.8 27B Alibaba · Apache 2.0 |
Image and video in, OSWorld 84 — the strongest local vision-language model at 24 GB. | Q4_K_M · 16.7 GB | 15 | Runs great |
| Best for agents | Qwen3.8 27B Alibaba · Apache 2.0 |
Computer use and tool calling are what 3.8 was trained for (OSWorld 84). | Q4_K_M · 16.7 GB | 15 | Runs great |
| Best translation | Qwen3.8 27B Alibaba · Apache 2.0 |
119 languages, and Qwen translations read as idiomatic rather than literal. | Q4_K_M · 16.7 GB | 15 | Runs great |
| Fastest good model | Qwen3.6 35B-A3B Alibaba · Apache 2.0 |
3.3B active per token: the fastest model that still competes with dense 27Bs. | Q4_K_M · 20.9 GB | 58 | Runs great |
| Best long context | Qwen3.8 27B Alibaba · Apache 2.0 |
262K native, and only 16 of 64 blocks keep a KV cache — long context is cheap here. | Q4_K_M · 24.2 GB at 128K context |
15 | Runs great |
Best local LLM for 128 GB unified memory
96.0 GB usable · 66 of 80 models fitThe 128 GB Mac or Strix Halo box: the 100B-class MoEs come into reach. Speeds below are estimated for a M4 Max · 128 GB; the fit verdicts hold for every 128 GB machine.
| Recommendation | Model | Why | Quant · total | Tok/s est. | |
|---|---|---|---|---|---|
| Best overall | Qwen3.8 27B Alibaba · Apache 2.0 |
Best capability per gigabyte: a current-generation dense 27B with vision and 262K context. | Q4_K_M · 16.7 GB | 20 | Runs great |
| Best for coding | Ling 3.0 Flash 124B-A5B inclusionAI · MIT |
SWE-bench Pro 56.6 claimed at 5B active — the 128 GB-class coding pick if you cannot fit DeepSeek V4. | Q4_K_M · 70.4 GB | 50 | Runs great |
| Best reasoning | Ling 3.0 Flash 124B-A5B inclusionAI · MIT |
AIME 93 claimed at 5B active; hybrid attention keeps long reasoning traces cheap. | Q4_K_M · 70.4 GB | 50 | Runs great |
| Best for writing | Llama 3.3 70B Instruct Meta · Llama 3.3 Community |
The creative-writing favourite: consistent voice, takes direction, handles fiction framing. | Q4_K_M · 42.8 GB | 7.7 | Runs great |
| Best vision | Qwen3.8 27B Alibaba · Apache 2.0 |
Image and video in, OSWorld 84 — the strongest local vision-language model at 24 GB. | Q4_K_M · 16.7 GB | 20 | Runs great |
| Best for agents | Ling 3.0 Flash 124B-A5B inclusionAI · MIT |
Agentic workflows are what Ling 3.0 was trained for; MIT, 5B active. | Q4_K_M · 70.4 GB | 50 | Runs great |
| Best translation | Qwen3.8 27B Alibaba · Apache 2.0 |
119 languages, and Qwen translations read as idiomatic rather than literal. | Q4_K_M · 16.7 GB | 20 | Runs great |
| Fastest good model | Qwen3.6 35B-A3B Alibaba · Apache 2.0 |
3.3B active per token: the fastest model that still competes with dense 27Bs. | Q4_K_M · 20.9 GB | 77 | Runs great |
| Best long context | Qwen3.8 27B Alibaba · Apache 2.0 |
262K native, and only 16 of 64 blocks keep a KV cache — long context is cheap here. | Q4_K_M · 24.2 GB at 128K context |
20 | Runs great |
By job
Best local LLM for coding, ranked per VRAM class with the whole shortlist, not just the pick.
By hardware
Best GPU for local LLMs — every card and Mac ranked by what it can run and how fast.
From scratch
How to run an LLM locally: pick hardware, pick a model, pick a runtime.