Best local LLM · 2026

The best local LLM for your VRAM, from 8 GB to 128 GB

Updated 8 Oct 2026 1 new this month

There is no single best local LLM — there is a best one for the memory you have. This page answers that for each VRAM class: one pick per job (general assistant, coding, reasoning, writing, vision, agents, translation, speed, long context), chosen from a hand-ordered shortlist and checked, by arithmetic, to fit at Q4_K_M with an 8K context. The ordering is editorial and says so; the sizes and speeds are computed, never benchmarked.

At a glance

Picks at Q4_K_M, 8K context, f16 KV cache. Not sure which class you are in? Detect your machine, or find it in the catalogue — every device page runs the same picks against your exact card.

Best local LLM for 8 GB VRAM

7.0 GB usable · 28 of 80 models fit

Entry-level and laptop cards. Small dense models and the first MoEs. Speeds below are estimated for a GeForce RTX 4060; the fit verdicts hold for every 8 GB card.

RecommendationModelWhyQuant · totalTok/s est.
Best overall Gemma 4 E4B
Google · Apache 2.0
Laptop class with images and audio in. Q4_K_M · 5.2 GB 37 Runs great
Best for coding Qwen3.5 4B
Alibaba · Apache 2.0
A coding agent in 3.4 GB — the 6–8 GB answer. Q4_K_M · 3.5 GB 63 Runs great
Best reasoning Qwen3.5 4B
Alibaba · Apache 2.0
Step-by-step reasoning in 8 GB. Q4_K_M · 3.5 GB 63 Runs great
Best for writing Gemma 4 E4B
Google · Apache 2.0
Readable prose on a laptop. Q4_K_M · 5.2 GB 37 Runs great
Best vision Gemma 4 E4B
Google · Apache 2.0
Laptop vision, audio too. Q4_K_M · 5.2 GB 37 Runs great
Best for agents Fara 7B
Microsoft · MIT
Web computer-use only — clicks and fills forms from screenshots, and stops to ask before anything irreversible. Q4_K_M · 5.7 GB 35 Runs great
Best translation Gemma 4 E4B
Google · Apache 2.0
140+ languages on a laptop. Q4_K_M · 5.2 GB 37 Runs great
Fastest good model LFM2.5 8B-A1B
Liquid AI · LFM Open License v1.0
1.5B active; designed for laptops without a GPU. Q4_K_M · 5.5 GB 75 Runs great
Best long context Qwen3.5 2B
Alibaba · Apache 2.0
Long documents on 4 GB. Q4_K_M · 3.4 GB
at 128K context
129 Runs great

Everything that fits 8 GB, in full →

Best local LLM for 12 GB VRAM

10.6 GB usable · 37 of 80 models fit

The most common upgrade step, and the first class where a 12B runs with headroom. Speeds below are estimated for a GeForce RTX 3060 12 GB; the fit verdicts hold for every 12 GB card.

RecommendationModelWhyQuant · totalTok/s est.
Best overall Gemma 4 12B
Google · Apache 2.0
The 12 GB generalist: text, image and audio in, 140+ languages. Q4_K_M · 8.2 GB 32 Runs great
Best for coding Qwen3.5 9B
Alibaba · Apache 2.0
The best coding model for 8–12 GB cards. Q4_K_M · 6.3 GB 40 Runs great
Best reasoning Qwen3.5 9B
Alibaba · Apache 2.0
Thinking mode on a 12 GB card — ahead of the 2025 R1 distills. Q4_K_M · 6.3 GB 40 Runs great
Best for writing Gemma 4 12B
Google · Apache 2.0
The 12 GB writing pick; 140+ languages. Q4_K_M · 8.2 GB 32 Runs great
Best vision Gemma 4 12B
Google · Apache 2.0
Image and audio understanding on 12 GB. Q4_K_M · 8.2 GB 32 Runs great
Best for agents Granite 4.1 8B
IBM · Apache 2.0
Enterprise-grade tool calling in a dense 8B, no thinking overhead. Q4_K_M · 6.8 GB 44 Runs great
Best translation Gemma 4 12B
Google · Apache 2.0
Broad language coverage on a 12 GB card. Q4_K_M · 8.2 GB 32 Runs great
Fastest good model LFM2.5 8B-A1B
Liquid AI · LFM Open License v1.0
1.5B active; designed for laptops without a GPU. Q4_K_M · 5.5 GB 99 Runs great
Best long context Qwen3.5 4B
Alibaba · Apache 2.0
262K on 8 GB. Q4_K_M · 7.2 GB
at 128K context
83 Runs great

Everything that fits 12 GB, in full →

Best local LLM for 16 GB VRAM

14.4 GB usable · 38 of 80 models fit

Where the 20B-class mixture-of-experts models start to fit whole. Speeds below are estimated for a GeForce RTX 5080; the fit verdicts hold for every 16 GB card.

RecommendationModelWhyQuant · totalTok/s est.
Best overall Gemma 4 12B
Google · Apache 2.0
The 12 GB generalist: text, image and audio in, 140+ languages. Q4_K_M · 8.2 GB 86 Runs great
Best for coding Qwen3.5 9B
Alibaba · Apache 2.0
The best coding model for 8–12 GB cards. Q4_K_M · 6.3 GB 107 Runs great
Best reasoning gpt-oss 20B
OpenAI · Apache 2.0
Low/medium/high reasoning effort in 13 GB. MXFP4 · 11.6 GB 120 Runs great
Best for writing Gemma 4 12B
Google · Apache 2.0
The 12 GB writing pick; 140+ languages. Q4_K_M · 8.2 GB 86 Runs great
Best vision Gemma 4 12B
Google · Apache 2.0
Image and audio understanding on 12 GB. Q4_K_M · 8.2 GB 86 Runs great
Best for agents gpt-oss 20B
OpenAI · Apache 2.0
Harmony-format tool calling, native 4-bit. MXFP4 · 11.6 GB 120 Runs great
Best translation Gemma 4 12B
Google · Apache 2.0
Broad language coverage on a 12 GB card. Q4_K_M · 8.2 GB 86 Runs great
Fastest good model gpt-oss 20B
OpenAI · Apache 2.0
3.6B active, native 4-bit. MXFP4 · 11.6 GB 120 Runs great
Best long context Qwen3.5 9B
Alibaba · Apache 2.0
262K native in 8–12 GB. Q4_K_M · 10.0 GB
at 128K context
107 Runs great

Everything that fits 16 GB, in full →

Best local LLM for 24 GB VRAM

22.4 GB usable · 58 of 80 models fit

The sweet spot: current-generation dense 27B models with room for context. Speeds below are estimated for a GeForce RTX 4090; the fit verdicts hold for every 24 GB card.

RecommendationModelWhyQuant · totalTok/s est.
Best overall Qwen3.8 27B
Alibaba · Apache 2.0
Best capability per gigabyte: a current-generation dense 27B with vision and 262K context. Q4_K_M · 16.7 GB 39 Runs great
Best for coding Qwen3.8 27B
Alibaba · Apache 2.0
Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. Q4_K_M · 16.7 GB 39 Runs great
Best reasoning Qwen3.8 27B
Alibaba · Apache 2.0
Thinking mode plus the best agentic reasoning at 24 GB. Q4_K_M · 16.7 GB 39 Runs great
Best for writing Gemma 4 26B-A4B
Google · Apache 2.0
Gemma 4 prose at MoE speed, with room for a long draft in context. Q4_K_M · 16.0 GB 110 Runs great
Best vision Qwen3.8 27B
Alibaba · Apache 2.0
Image and video in, OSWorld 84 — the strongest local vision-language model at 24 GB. Q4_K_M · 16.7 GB 39 Runs great
Best for agents Qwen3.8 27B
Alibaba · Apache 2.0
Computer use and tool calling are what 3.8 was trained for (OSWorld 84). Q4_K_M · 16.7 GB 39 Runs great
Best translation Qwen3.8 27B
Alibaba · Apache 2.0
119 languages, and Qwen translations read as idiomatic rather than literal. Q4_K_M · 16.7 GB 39 Runs great
Fastest good model Gemma 4 26B-A4B
Google · Apache 2.0
3.8B active and a small file — the quickest Gemma 4 with vision. Q4_K_M · 16.0 GB 110 Runs great
Best long context Nemotron 3.5 Lightning 30B-A3B
NVIDIA · OpenMDW-1.1
262K, and a Mamba hybrid whose cache barely grows. Q4_K_M · 19.1 GB
at 128K context
130 Tight fit

Everything that fits 24 GB, in full →

Best local LLM for 32 GB VRAM

30.4 GB usable · 58 of 80 models fit

Dense 30Bs at Q6, or a 27B with a very long context. Speeds below are estimated for a GeForce RTX 5090; the fit verdicts hold for every 32 GB card.

RecommendationModelWhyQuant · totalTok/s est.
Best overall Qwen3.8 27B
Alibaba · Apache 2.0
Best capability per gigabyte: a current-generation dense 27B with vision and 262K context. Q4_K_M · 16.7 GB 69 Runs great
Best for coding Qwen3.8 27B
Alibaba · Apache 2.0
Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. Q4_K_M · 16.7 GB 69 Runs great
Best reasoning Gemma 4 31B
Google · Apache 2.0
89% AIME — the strongest maths of any model under 32B. Q4_K_M · 20.2 GB 62 Runs great
Best for writing Gemma 4 31B
Google · Apache 2.0
The cleanest, least verbose prose of the 2026 24–32 GB class. Q4_K_M · 20.2 GB 62 Runs great
Best vision Qwen3.8 27B
Alibaba · Apache 2.0
Image and video in, OSWorld 84 — the strongest local vision-language model at 24 GB. Q4_K_M · 16.7 GB 69 Runs great
Best for agents Qwen3.8 27B
Alibaba · Apache 2.0
Computer use and tool calling are what 3.8 was trained for (OSWorld 84). Q4_K_M · 16.7 GB 69 Runs great
Best translation Qwen3.8 27B
Alibaba · Apache 2.0
119 languages, and Qwen translations read as idiomatic rather than literal. Q4_K_M · 16.7 GB 69 Runs great
Fastest good model Qwen3.6 35B-A3B
Alibaba · Apache 2.0
3.3B active per token: the fastest model that still competes with dense 27Bs. Q4_K_M · 20.9 GB 225 Runs great
Best long context Qwen3.8 27B
Alibaba · Apache 2.0
262K native, and only 16 of 64 blocks keep a KV cache — long context is cheap here. Q4_K_M · 24.2 GB
at 128K context
69 Runs great
Cards in this class RTX 5090RTX 5000 AdaPro W7800

Everything that fits 32 GB, in full →

Best local LLM for 48 GB VRAM

46.4 GB usable · 60 of 80 models fit

Workstation class: 70B at Q4, or a 30B near-lossless. Speeds below are estimated for a RTX 6000 Ada; the fit verdicts hold for every 48 GB card.

RecommendationModelWhyQuant · totalTok/s est.
Best overall Qwen3.8 27B
Alibaba · Apache 2.0
Best capability per gigabyte: a current-generation dense 27B with vision and 262K context. Q4_K_M · 16.7 GB 37 Runs great
Best for coding Qwen3.8 27B
Alibaba · Apache 2.0
Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. Q4_K_M · 16.7 GB 37 Runs great
Best reasoning Gemma 4 31B
Google · Apache 2.0
89% AIME — the strongest maths of any model under 32B. Q4_K_M · 20.2 GB 33 Runs great
Best for writing Gemma 4 31B
Google · Apache 2.0
The cleanest, least verbose prose of the 2026 24–32 GB class. Q4_K_M · 20.2 GB 33 Runs great
Best vision Qwen3.8 27B
Alibaba · Apache 2.0
Image and video in, OSWorld 84 — the strongest local vision-language model at 24 GB. Q4_K_M · 16.7 GB 37 Runs great
Best for agents Qwen3.8 27B
Alibaba · Apache 2.0
Computer use and tool calling are what 3.8 was trained for (OSWorld 84). Q4_K_M · 16.7 GB 37 Runs great
Best translation Qwen3.8 27B
Alibaba · Apache 2.0
119 languages, and Qwen translations read as idiomatic rather than literal. Q4_K_M · 16.7 GB 37 Runs great
Fastest good model Qwen3.6 35B-A3B
Alibaba · Apache 2.0
3.3B active per token: the fastest model that still competes with dense 27Bs. Q4_K_M · 20.9 GB 120 Runs great
Best long context Qwen3.8 27B
Alibaba · Apache 2.0
262K native, and only 16 of 64 blocks keep a KV cache — long context is cheap here. Q4_K_M · 24.2 GB
at 128K context
37 Runs great

Everything that fits 48 GB, in full →

Best local LLM for 64 GB unified memory

48.0 GB usable · 60 of 80 models fit

The 64 GB Mac. Unified memory, so the GPU can address about three quarters of it. Speeds below are estimated for a M4 Max · 64 GB; the fit verdicts hold for every 64 GB machine.

RecommendationModelWhyQuant · totalTok/s est.
Best overall Qwen3.8 27B
Alibaba · Apache 2.0
Best capability per gigabyte: a current-generation dense 27B with vision and 262K context. Q4_K_M · 16.7 GB 15 Runs great
Best for coding Qwen3.8 27B
Alibaba · Apache 2.0
Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. Q4_K_M · 16.7 GB 15 Runs great
Best reasoning Gemma 4 31B
Google · Apache 2.0
89% AIME — the strongest maths of any model under 32B. Q4_K_M · 20.2 GB 13 Runs great
Best for writing Gemma 4 31B
Google · Apache 2.0
The cleanest, least verbose prose of the 2026 24–32 GB class. Q4_K_M · 20.2 GB 13 Runs great
Best vision Qwen3.8 27B
Alibaba · Apache 2.0
Image and video in, OSWorld 84 — the strongest local vision-language model at 24 GB. Q4_K_M · 16.7 GB 15 Runs great
Best for agents Qwen3.8 27B
Alibaba · Apache 2.0
Computer use and tool calling are what 3.8 was trained for (OSWorld 84). Q4_K_M · 16.7 GB 15 Runs great
Best translation Qwen3.8 27B
Alibaba · Apache 2.0
119 languages, and Qwen translations read as idiomatic rather than literal. Q4_K_M · 16.7 GB 15 Runs great
Fastest good model Qwen3.6 35B-A3B
Alibaba · Apache 2.0
3.3B active per token: the fastest model that still competes with dense 27Bs. Q4_K_M · 20.9 GB 58 Runs great
Best long context Qwen3.8 27B
Alibaba · Apache 2.0
262K native, and only 16 of 64 blocks keep a KV cache — long context is cheap here. Q4_K_M · 24.2 GB
at 128K context
15 Runs great

Every model against the M4 Max · 64 GB →

Best local LLM for 128 GB unified memory

96.0 GB usable · 66 of 80 models fit

The 128 GB Mac or Strix Halo box: the 100B-class MoEs come into reach. Speeds below are estimated for a M4 Max · 128 GB; the fit verdicts hold for every 128 GB machine.

RecommendationModelWhyQuant · totalTok/s est.
Best overall Qwen3.8 27B
Alibaba · Apache 2.0
Best capability per gigabyte: a current-generation dense 27B with vision and 262K context. Q4_K_M · 16.7 GB 20 Runs great
Best for coding Ling 3.0 Flash 124B-A5B
inclusionAI · MIT
SWE-bench Pro 56.6 claimed at 5B active — the 128 GB-class coding pick if you cannot fit DeepSeek V4. Q4_K_M · 70.4 GB 50 Runs great
Best reasoning Ling 3.0 Flash 124B-A5B
inclusionAI · MIT
AIME 93 claimed at 5B active; hybrid attention keeps long reasoning traces cheap. Q4_K_M · 70.4 GB 50 Runs great
Best for writing Llama 3.3 70B Instruct
Meta · Llama 3.3 Community
The creative-writing favourite: consistent voice, takes direction, handles fiction framing. Q4_K_M · 42.8 GB 7.7 Runs great
Best vision Qwen3.8 27B
Alibaba · Apache 2.0
Image and video in, OSWorld 84 — the strongest local vision-language model at 24 GB. Q4_K_M · 16.7 GB 20 Runs great
Best for agents Ling 3.0 Flash 124B-A5B
inclusionAI · MIT
Agentic workflows are what Ling 3.0 was trained for; MIT, 5B active. Q4_K_M · 70.4 GB 50 Runs great
Best translation Qwen3.8 27B
Alibaba · Apache 2.0
119 languages, and Qwen translations read as idiomatic rather than literal. Q4_K_M · 16.7 GB 20 Runs great
Fastest good model Qwen3.6 35B-A3B
Alibaba · Apache 2.0
3.3B active per token: the fastest model that still competes with dense 27Bs. Q4_K_M · 20.9 GB 77 Runs great
Best long context Qwen3.8 27B
Alibaba · Apache 2.0
262K native, and only 16 of 64 blocks keep a KV cache — long context is cheap here. Q4_K_M · 24.2 GB
at 128K context
20 Runs great

Every model against the M4 Max · 128 GB →


By job

Best local LLM for coding, ranked per VRAM class with the whole shortlist, not just the pick.

By hardware

Best GPU for local LLMs — every card and Mac ranked by what it can run and how fast.

From scratch

How to run an LLM locally: pick hardware, pick a model, pick a runtime.