Best local LLM for coding · 2026
The best local LLM for coding, ranked for your VRAM
The coding shortlist is ordered by what the vendors' own evaluations and the local-model community report for software-engineering agents, code editing and autocomplete — SWE-bench Verified, Terminal-Bench, DeepSWE — as of the catalogue date. For each VRAM class below we list every model on that list that actually fits at Q4_K_M with an 8K context, so the top row is the pick and the rest are what you give up by choosing it. Numbers are computed, not benchmarked; how the shortlist works.
A coding agent needs context more than most jobs. The "max context" column is the longest window the model can hold on this class at that quantisation — if it reads under 32K, expect to drop a quantisation step or run the model with a q8_0 KV cache before pointing an agent at a whole repository.
Best coding LLM for 8 GB VRAM
7.0 GB usable · speeds for a GeForce RTX 4060Qwen3.5 9B is the pick: The best coding model for 8–12 GB cards. 4 other models on the coding shortlist also fit, in order of capability:
| # | Model | Why | Params | Quant · total | Max context | Tok/s est. | |
|---|---|---|---|---|---|---|---|
| 1 | Qwen3.5 9B Alibaba · Apache 2.0 |
The best coding model for 8–12 GB cards. | 9.65B | Q4_K_M · 6.3 GB | 31K | 30 | Tight fit |
| 2 | Ornith 1.5 9B Ornith AI · MIT |
A coding-agent reasoning build on the Qwen3.5 9B architecture. | 9.41B | Q4_K_M · 6.1 GB | 35K | 31 | Tight fit |
| 3 | Qwen3.5 4B Alibaba · Apache 2.0 |
A coding agent in 3.4 GB — the 6–8 GB answer. | 4.66B | Q4_K_M · 3.5 GB | 120K | 63 | Runs great |
| 4 | Qwen2.5-Coder 7B Alibaba · Apache 2.0 |
Still the best fill-in-the-middle autocomplete model under 8B. | 7.62B | Q4_K_M · 5.3 GB | 38K | 38 | Runs great |
| 5 | Qwen3.5 2B Alibaba · Apache 2.0 |
Autocomplete on 4 GB. | 2.27B | Q4_K_M · 2.0 GB | 256K | 129 | Runs great |
Best coding LLM for 12 GB VRAM
10.6 GB usable · speeds for a GeForce RTX 3060 12 GBQwen3.5 9B is the pick: The best coding model for 8–12 GB cards. 4 other models on the coding shortlist also fit, in order of capability:
| # | Model | Why | Params | Quant · total | Max context | Tok/s est. | |
|---|---|---|---|---|---|---|---|
| 1 | Qwen3.5 9B Alibaba · Apache 2.0 |
The best coding model for 8–12 GB cards. | 9.65B | Q4_K_M · 6.3 GB | 146K | 40 | Runs great |
| 2 | Ornith 1.5 9B Ornith AI · MIT |
A coding-agent reasoning build on the Qwen3.5 9B architecture. | 9.41B | Q4_K_M · 6.1 GB | 150K | 41 | Runs great |
| 3 | Qwen3.5 4B Alibaba · Apache 2.0 |
A coding agent in 3.4 GB — the 6–8 GB answer. | 4.66B | Q4_K_M · 3.5 GB | 236K | 83 | Runs great |
| 4 | Qwen2.5-Coder 7B Alibaba · Apache 2.0 |
Still the best fill-in-the-middle autocomplete model under 8B. | 7.62B | Q4_K_M · 5.3 GB | 104K | 51 | Runs great |
| 5 | Qwen3.5 2B Alibaba · Apache 2.0 |
Autocomplete on 4 GB. | 2.27B | Q4_K_M · 2.0 GB | 256K | 171 | Runs great |
Best coding LLM for 16 GB VRAM
14.4 GB usable · speeds for a GeForce RTX 5080Qwen3.5 9B is the pick: The best coding model for 8–12 GB cards. 4 other models on the coding shortlist also fit, in order of capability:
| # | Model | Why | Params | Quant · total | Max context | Tok/s est. | |
|---|---|---|---|---|---|---|---|
| 1 | Qwen3.5 9B Alibaba · Apache 2.0 |
The best coding model for 8–12 GB cards. | 9.65B | Q4_K_M · 6.3 GB | 256K | 107 | Runs great |
| 2 | Ornith 1.5 9B Ornith AI · MIT |
A coding-agent reasoning build on the Qwen3.5 9B architecture. | 9.41B | Q4_K_M · 6.1 GB | 256K | 110 | Runs great |
| 3 | Qwen3.5 4B Alibaba · Apache 2.0 |
A coding agent in 3.4 GB — the 6–8 GB answer. | 4.66B | Q4_K_M · 3.5 GB | 256K | 222 | Runs great |
| 4 | Qwen2.5-Coder 7B Alibaba · Apache 2.0 |
Still the best fill-in-the-middle autocomplete model under 8B. | 7.62B | Q4_K_M · 5.3 GB | 128K | 136 | Runs great |
| 5 | Qwen3.5 2B Alibaba · Apache 2.0 |
Autocomplete on 4 GB. | 2.27B | Q4_K_M · 2.0 GB | 256K | 455 | Runs great |
Best coding LLM for 24 GB VRAM
22.4 GB usable · speeds for a GeForce RTX 4090Qwen3.8 27B is the pick: Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. 11 other models on the coding shortlist also fit, in order of capability:
| # | Model | Why | Params | Quant · total | Max context | Tok/s est. | |
|---|---|---|---|---|---|---|---|
| 1 | Qwen3.8 27B Alibaba · Apache 2.0 |
Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. | 27.8B | Q4_K_M · 16.7 GB | 98K | 39 | Runs great |
| 2 | Qwen3.6 27B Alibaba · Apache 2.0 |
77.2 SWE-bench Verified at 27B; the spring-2026 local coding standard. | 27.8B | Q4_K_M · 16.7 GB | 98K | 39 | Runs great |
| 3 | Qwen3.6 35B-A3B Alibaba · Apache 2.0 |
Agentic coding at 3B-active speed — the best MoE under 40B for it. | 35.9B MoE | Q4_K_M · 20.9 GB | 82K | 126 | Tight fit |
| 4 | Devstral Small 2 24B Mistral AI · Apache 2.0 |
Tuned for software-engineering agents like OpenHands and Cline. | 24B | Q4_K_M · 15.3 GB | 53K | 45 | Runs great |
| 5 | GLM-4.7-Flash 30B-A3B Z.ai · MIT |
MIT-licensed 30B-A3B built for agentic coding; 60–80 tok/s reported on a 4090. | 31.2B MoE | Q4_K_M · 18.6 GB | 82K | 139 | Runs great |
| 6 | Ornith 1.5 35B-A3B Ornith AI · MIT |
Reasoning-first coding build on the Qwen3.6 MoE; thinks before it edits. | 35.9B MoE | Q4_K_M · 20.9 GB | 82K | 126 | Tight fit |
| 7 | Qwen3 Coder 30B-A3B Alibaba · Apache 2.0 |
The most-downloaded local code model; 256K window. | 30.5B MoE | Q4_K_M · 18.5 GB | 49K | 126 | Runs great |
| 8 | Qwen3.5 9B Alibaba · Apache 2.0 |
The best coding model for 8–12 GB cards. | 9.65B | Q4_K_M · 6.3 GB | 256K | 112 | Runs great |
| 9 | Ornith 1.5 9B Ornith AI · MIT |
A coding-agent reasoning build on the Qwen3.5 9B architecture. | 9.41B | Q4_K_M · 6.1 GB | 256K | 115 | Runs great |
| 10 | Qwen3.5 4B Alibaba · Apache 2.0 |
A coding agent in 3.4 GB — the 6–8 GB answer. | 4.66B | Q4_K_M · 3.5 GB | 256K | 233 | Runs great |
| 11 | Qwen2.5-Coder 7B Alibaba · Apache 2.0 |
Still the best fill-in-the-middle autocomplete model under 8B. | 7.62B | Q4_K_M · 5.3 GB | 128K | 142 | Runs great |
| 12 | Qwen3.5 2B Alibaba · Apache 2.0 |
Autocomplete on 4 GB. | 2.27B | Q4_K_M · 2.0 GB | 256K | 478 | Runs great |
Best coding LLM for 32 GB VRAM
30.4 GB usable · speeds for a GeForce RTX 5090Qwen3.8 27B is the pick: Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. 11 other models on the coding shortlist also fit, in order of capability:
| # | Model | Why | Params | Quant · total | Max context | Tok/s est. | |
|---|---|---|---|---|---|---|---|
| 1 | Qwen3.8 27B Alibaba · Apache 2.0 |
Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. | 27.8B | Q4_K_M · 16.7 GB | 226K | 69 | Runs great |
| 2 | Qwen3.6 27B Alibaba · Apache 2.0 |
77.2 SWE-bench Verified at 27B; the spring-2026 local coding standard. | 27.8B | Q4_K_M · 16.7 GB | 226K | 69 | Runs great |
| 3 | Qwen3.6 35B-A3B Alibaba · Apache 2.0 |
Agentic coding at 3B-active speed — the best MoE under 40B for it. | 35.9B MoE | Q4_K_M · 20.9 GB | 256K | 225 | Runs great |
| 4 | Devstral Small 2 24B Mistral AI · Apache 2.0 |
Tuned for software-engineering agents like OpenHands and Cline. | 24B | Q4_K_M · 15.3 GB | 104K | 80 | Runs great |
| 5 | GLM-4.7-Flash 30B-A3B Z.ai · MIT |
MIT-licensed 30B-A3B built for agentic coding; 60–80 tok/s reported on a 4090. | 31.2B MoE | Q4_K_M · 18.6 GB | 198K | 247 | Runs great |
| 6 | Ornith 1.5 35B-A3B Ornith AI · MIT |
Reasoning-first coding build on the Qwen3.6 MoE; thinks before it edits. | 35.9B MoE | Q4_K_M · 20.9 GB | 256K | 225 | Runs great |
| 7 | Qwen3 Coder 30B-A3B Alibaba · Apache 2.0 |
The most-downloaded local code model; 256K window. | 30.5B MoE | Q4_K_M · 18.5 GB | 134K | 225 | Runs great |
| 8 | Qwen3.5 9B Alibaba · Apache 2.0 |
The best coding model for 8–12 GB cards. | 9.65B | Q4_K_M · 6.3 GB | 256K | 200 | Runs great |
| 9 | Ornith 1.5 9B Ornith AI · MIT |
A coding-agent reasoning build on the Qwen3.5 9B architecture. | 9.41B | Q4_K_M · 6.1 GB | 256K | 205 | Runs great |
| 10 | Qwen3.5 4B Alibaba · Apache 2.0 |
A coding agent in 3.4 GB — the 6–8 GB answer. | 4.66B | Q4_K_M · 3.5 GB | 256K | 414 | Runs great |
| 11 | Qwen2.5-Coder 7B Alibaba · Apache 2.0 |
Still the best fill-in-the-middle autocomplete model under 8B. | 7.62B | Q4_K_M · 5.3 GB | 128K | 253 | Runs great |
| 12 | Qwen3.5 2B Alibaba · Apache 2.0 |
Autocomplete on 4 GB. | 2.27B | Q4_K_M · 2.0 GB | 256K | 850 | Runs great |
Best coding LLM for 48 GB VRAM
46.4 GB usable · speeds for a RTX 6000 AdaQwen3.8 27B is the pick: Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. 11 other models on the coding shortlist also fit, in order of capability:
| # | Model | Why | Params | Quant · total | Max context | Tok/s est. | |
|---|---|---|---|---|---|---|---|
| 1 | Qwen3.8 27B Alibaba · Apache 2.0 |
Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. | 27.8B | Q4_K_M · 16.7 GB | 256K | 37 | Runs great |
| 2 | Qwen3.6 27B Alibaba · Apache 2.0 |
77.2 SWE-bench Verified at 27B; the spring-2026 local coding standard. | 27.8B | Q4_K_M · 16.7 GB | 256K | 37 | Runs great |
| 3 | Qwen3.6 35B-A3B Alibaba · Apache 2.0 |
Agentic coding at 3B-active speed — the best MoE under 40B for it. | 35.9B MoE | Q4_K_M · 20.9 GB | 256K | 120 | Runs great |
| 4 | Devstral Small 2 24B Mistral AI · Apache 2.0 |
Tuned for software-engineering agents like OpenHands and Cline. | 24B | Q4_K_M · 15.3 GB | 206K | 43 | Runs great |
| 5 | GLM-4.7-Flash 30B-A3B Z.ai · MIT |
MIT-licensed 30B-A3B built for agentic coding; 60–80 tok/s reported on a 4090. | 31.2B MoE | Q4_K_M · 18.6 GB | 198K | 133 | Runs great |
| 6 | Ornith 1.5 35B-A3B Ornith AI · MIT |
Reasoning-first coding build on the Qwen3.6 MoE; thinks before it edits. | 35.9B MoE | Q4_K_M · 20.9 GB | 256K | 120 | Runs great |
| 7 | Qwen3 Coder 30B-A3B Alibaba · Apache 2.0 |
The most-downloaded local code model; 256K window. | 30.5B MoE | Q4_K_M · 18.5 GB | 256K | 120 | Runs great |
| 8 | Qwen3.5 9B Alibaba · Apache 2.0 |
The best coding model for 8–12 GB cards. | 9.65B | Q4_K_M · 6.3 GB | 256K | 107 | Runs great |
| 9 | Ornith 1.5 9B Ornith AI · MIT |
A coding-agent reasoning build on the Qwen3.5 9B architecture. | 9.41B | Q4_K_M · 6.1 GB | 256K | 110 | Runs great |
| 10 | Qwen3.5 4B Alibaba · Apache 2.0 |
A coding agent in 3.4 GB — the 6–8 GB answer. | 4.66B | Q4_K_M · 3.5 GB | 256K | 222 | Runs great |
| 11 | Qwen2.5-Coder 7B Alibaba · Apache 2.0 |
Still the best fill-in-the-middle autocomplete model under 8B. | 7.62B | Q4_K_M · 5.3 GB | 128K | 136 | Runs great |
| 12 | Qwen3.5 2B Alibaba · Apache 2.0 |
Autocomplete on 4 GB. | 2.27B | Q4_K_M · 2.0 GB | 256K | 455 | Runs great |
Best coding LLM for 64 GB unified memory
48.0 GB usable · speeds for a M4 Max · 64 GBQwen3.8 27B is the pick: Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. 11 other models on the coding shortlist also fit, in order of capability:
| # | Model | Why | Params | Quant · total | Max context | Tok/s est. | |
|---|---|---|---|---|---|---|---|
| 1 | Qwen3.8 27B Alibaba · Apache 2.0 |
Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. | 27.8B | Q4_K_M · 16.7 GB | 256K | 15 | Runs great |
| 2 | Qwen3.6 27B Alibaba · Apache 2.0 |
77.2 SWE-bench Verified at 27B; the spring-2026 local coding standard. | 27.8B | Q4_K_M · 16.7 GB | 256K | 15 | Runs great |
| 3 | Qwen3.6 35B-A3B Alibaba · Apache 2.0 |
Agentic coding at 3B-active speed — the best MoE under 40B for it. | 35.9B MoE | Q4_K_M · 20.9 GB | 256K | 58 | Runs great |
| 4 | Devstral Small 2 24B Mistral AI · Apache 2.0 |
Tuned for software-engineering agents like OpenHands and Cline. | 24B | Q4_K_M · 15.3 GB | 216K | 17 | Runs great |
| 5 | GLM-4.7-Flash 30B-A3B Z.ai · MIT |
MIT-licensed 30B-A3B built for agentic coding; 60–80 tok/s reported on a 4090. | 31.2B MoE | Q4_K_M · 18.6 GB | 198K | 63 | Runs great |
| 6 | Ornith 1.5 35B-A3B Ornith AI · MIT |
Reasoning-first coding build on the Qwen3.6 MoE; thinks before it edits. | 35.9B MoE | Q4_K_M · 20.9 GB | 256K | 58 | Runs great |
| 7 | Qwen3 Coder 30B-A3B Alibaba · Apache 2.0 |
The most-downloaded local code model; 256K window. | 30.5B MoE | Q4_K_M · 18.5 GB | 256K | 58 | Runs great |
| 8 | Qwen3.5 9B Alibaba · Apache 2.0 |
The best coding model for 8–12 GB cards. | 9.65B | Q4_K_M · 6.3 GB | 256K | 42 | Runs great |
| 9 | Ornith 1.5 9B Ornith AI · MIT |
A coding-agent reasoning build on the Qwen3.5 9B architecture. | 9.41B | Q4_K_M · 6.1 GB | 256K | 43 | Runs great |
| 10 | Qwen3.5 4B Alibaba · Apache 2.0 |
A coding agent in 3.4 GB — the 6–8 GB answer. | 4.66B | Q4_K_M · 3.5 GB | 256K | 87 | Runs great |
| 11 | Qwen2.5-Coder 7B Alibaba · Apache 2.0 |
Still the best fill-in-the-middle autocomplete model under 8B. | 7.62B | Q4_K_M · 5.3 GB | 128K | 53 | Runs great |
| 12 | Qwen3.5 2B Alibaba · Apache 2.0 |
Autocomplete on 4 GB. | 2.27B | Q4_K_M · 2.0 GB | 256K | 179 | Runs great |
Best coding LLM for 128 GB unified memory
96.0 GB usable · speeds for a M4 Max · 128 GBLing 3.0 Flash 124B-A5B is the pick: SWE-bench Pro 56.6 claimed at 5B active — the 128 GB-class coding pick if you cannot fit DeepSeek V4. 12 other models on the coding shortlist also fit, in order of capability:
| # | Model | Why | Params | Quant · total | Max context | Tok/s est. | |
|---|---|---|---|---|---|---|---|
| 1 | Ling 3.0 Flash 124B-A5B inclusionAI · MIT |
SWE-bench Pro 56.6 claimed at 5B active — the 128 GB-class coding pick if you cannot fit DeepSeek V4. | 124B MoE | Q4_K_M · 70.4 GB | 256K | 50 | Runs great |
| 2 | Qwen3.8 27B Alibaba · Apache 2.0 |
Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. | 27.8B | Q4_K_M · 16.7 GB | 256K | 20 | Runs great |
| 3 | Qwen3.6 27B Alibaba · Apache 2.0 |
77.2 SWE-bench Verified at 27B; the spring-2026 local coding standard. | 27.8B | Q4_K_M · 16.7 GB | 256K | 20 | Runs great |
| 4 | Qwen3.6 35B-A3B Alibaba · Apache 2.0 |
Agentic coding at 3B-active speed — the best MoE under 40B for it. | 35.9B MoE | Q4_K_M · 20.9 GB | 256K | 77 | Runs great |
| 5 | Devstral Small 2 24B Mistral AI · Apache 2.0 |
Tuned for software-engineering agents like OpenHands and Cline. | 24B | Q4_K_M · 15.3 GB | 384K | 23 | Runs great |
| 6 | GLM-4.7-Flash 30B-A3B Z.ai · MIT |
MIT-licensed 30B-A3B built for agentic coding; 60–80 tok/s reported on a 4090. | 31.2B MoE | Q4_K_M · 18.6 GB | 198K | 84 | Runs great |
| 7 | Ornith 1.5 35B-A3B Ornith AI · MIT |
Reasoning-first coding build on the Qwen3.6 MoE; thinks before it edits. | 35.9B MoE | Q4_K_M · 20.9 GB | 256K | 77 | Runs great |
| 8 | Qwen3 Coder 30B-A3B Alibaba · Apache 2.0 |
The most-downloaded local code model; 256K window. | 30.5B MoE | Q4_K_M · 18.5 GB | 256K | 77 | Runs great |
| 9 | Qwen3.5 9B Alibaba · Apache 2.0 |
The best coding model for 8–12 GB cards. | 9.65B | Q4_K_M · 6.3 GB | 256K | 56 | Runs great |
| 10 | Ornith 1.5 9B Ornith AI · MIT |
A coding-agent reasoning build on the Qwen3.5 9B architecture. | 9.41B | Q4_K_M · 6.1 GB | 256K | 58 | Runs great |
| 11 | Qwen3.5 4B Alibaba · Apache 2.0 |
A coding agent in 3.4 GB — the 6–8 GB answer. | 4.66B | Q4_K_M · 3.5 GB | 256K | 116 | Runs great |
| 12 | Qwen2.5-Coder 7B Alibaba · Apache 2.0 |
Still the best fill-in-the-middle autocomplete model under 8B. | 7.62B | Q4_K_M · 5.3 GB | 128K | 71 | Runs great |
| 13 | Qwen3.5 2B Alibaba · Apache 2.0 |
Autocomplete on 4 GB. | 2.27B | Q4_K_M · 2.0 GB | 256K | 239 | Runs great |
The whole coding shortlist, and what each one needs
Every model on the list, best first, with the smallest discrete card that runs it with 15% headroom at the recommended quantisation and 8K context.
| # | Model | Params | Weights | Needs | Ollama tag |
|---|---|---|---|---|---|
| 1 | DeepSeek V4.1 Flash 552B-A16B New Twice V4 Flash's backbone (552B, 16B active per generated token) plus 196B of Engram lookup tables (189 GiB at FP8) that DwarfStar streams from SSD, so they are not counted here. A 512 GB Mac model, and as of October 2026 no mainline llama.cpp or Ollama build: DwarfStar on a Mac, vLLM across four GPUs. DeepSeek puts the cache at 890 bytes/token; modelled conservatively as 4 latent layers plus a 128-token window. |
552B MoE | 310.4 GB @ Q4_K_M | more than one card | — |
| 2 | Kimi K2.6 1T-A32B The open coding-agent benchmark leader of spring 2026 (80.2 SWE-bench). A 512 GB Mac Studio pair, or a ceiling. |
1027B MoE | 577.5 GB @ Q4_K_M | more than one card | — |
| 3 | GLM-5.3 744B-A40B Same base as GLM-5.2, new post-training: Z.ai's most capable open-weights coder (Terminal-Bench 3.0 28.3, DeepSWE 66.9). Not MIT any more — its own GLM-5.3 licence. Same ceiling as 5.2: 512 GB of unified memory at Q4. |
753B MoE | 423.4 GB @ Q4_K_M | more than one card | — |
| 4 | GLM-5.3-Flash 320B-A18B The first natively multimodal GLM-5 and the first hybrid: 34 linear-attention blocks and 11 sparse-attention blocks with a 512-wide latent cache, so a 1M window stays affordable. Z.ai says it beats GLM-5.2 at 18B active; MIT. |
321B MoE | 180.5 GB @ Q4_K_M | more than one card | — |
| 5 | DeepSeek V4 Flash 284B-A13B The V4 for 192–256 GB machines: Q3 squeezes into 192 GB, Q4 wants 256 GB, and 128 GB falls short even at Q2. Cache is modelled as a 576-wide latent; V4 compresses it further at long context, so this is conservative. |
284B MoE | 159.7 GB @ Q4_K_M | 192 GB card | — |
| 6 | Qwen3.8-Flash-Next 180B-A6B The open preview of the Qwen4 architecture: a 125B-A6B hybrid (Gated DeltaNet + sparse attention, KV cache on 12 of 48 blocks) plus a 51B n-gram embedding and a 4B draft head — 180B on disk, 6B active. The hosted "Qwen3.8-Flash" is this model with a 1M window. |
180B MoE | 101.2 GB @ Q4_K_M | 141 GB card | qwen3.8-flash-next:125b-a6b-nvfp4 |
| 7 | Ling 3.0 Flash 124B-A5B A 124B hybrid (5 linear-attention layers per MLA layer) with 5.1B active: SWE-bench Pro 56.6 and AIME 93 claimed. Built for 96–128 GB machines. |
124B MoE | 69.7 GB @ Q4_K_M | 141 GB card | — |
| 8 | Qwen3.8 27B The current default local Qwen: dense 27B, text + image + video, 262K context. Only 16 of its 64 blocks keep a KV cache, so long context is cheap. |
27.8B | 15.6 GB @ Q4_K_M | 24 GB card | qwen3.8:27b |
| 9 | Qwen3.6 27B The 24 GB coding pick of spring 2026 (77.2 SWE-bench Verified). Same shape as 3.8, one generation behind. |
27.8B | 15.6 GB @ Q4_K_M | 24 GB card | qwen3.6:27b |
| 10 | Qwen3.6 35B-A3B Mixture of experts with ~3B active: the fastest serious model a 24 GB card runs, and the best MoE under 40B on agentic coding. |
35.9B MoE | 20.2 GB @ Q4_K_M | 32 GB card | qwen3.6:35b-a3b |
| 11 | Devstral Small 2 24B Built for software-engineering agents (OpenHands, Cline). Dense 24B, 384K window. Not in the Ollama library. |
24B | 13.5 GB @ Q4_K_M | 20 GB card | — |
| 12 | GLM-4.7-Flash 30B-A3B MIT-licensed 30B-A3B tuned for agentic coding, with a DeepSeek-style latent KV cache. 60–80 tok/s reported on a 4090. |
31.2B MoE | 17.5 GB @ Q4_K_M | 24 GB card | glm-4.7-flash:latest |
| 13 | Ornith 1.5 35B-A3B A reasoning-first MIT build on the Qwen3.6 35B-A3B architecture (thinks before every answer). Same VRAM as its base. |
35.9B MoE | 20.2 GB @ Q4_K_M | 32 GB card | ornith-1.5:35b |
| 14 | Qwen3 Coder 30B-A3B Agentic coding MoE with a 256K native window. Still the most-downloaded local code model. |
30.5B MoE | 17.1 GB @ Q4_K_M | 24 GB card | qwen3-coder:30b |
| 15 | Qwen3.5 9B The default for 8–12 GB cards in 2026: beats every older 8B on every published benchmark, with vision. |
9.65B | 5.4 GB @ Q4_K_M | 10 GB card | qwen3.5:9b |
| 16 | Ornith 1.5 9B The small Ornith: a coding-agent reasoning build on the Qwen3.5 9B architecture. Same VRAM as its base, thinks before every answer. |
9.41B | 5.3 GB @ Q4_K_M | 10 GB card | ornith-1.5:9b |
| 17 | Qwen3.5 4B The 8 GB coding agent. Q4 lands near 3.4 GB, leaving room for a real context window. |
4.66B | 2.6 GB @ Q4_K_M | 8 GB card | qwen3.5:4b |
| 18 | Qwen2.5-Coder 7B The standard local autocomplete model — small enough to keep resident all day, and still the best FIM model under 8B. |
7.62B | 4.3 GB @ Q4_K_M | 8 GB card | qwen2.5-coder:7b |
| 19 | Qwen3.5 2B Phone-class, and multimodal. Replaces Llama 3.2 3B as the "it runs on anything" answer. |
2.27B | 1.3 GB @ Q4_K_M | 8 GB card | qwen3.5:2b |
Autocomplete and agents want different things. Fill-in-the-middle completion is latency-bound and happy with a 4–9B model; an agent that edits files and runs tests wants the strongest model you can hold with a 32K+ window, and will wait for it. If you are unsure which card to buy for this, the GPU ranking says what each class unlocks.