Best local LLM for coding · 2026

The best local LLM for coding, ranked for your VRAM

Updated 8 Oct 2026 1 new this month

The coding shortlist is ordered by what the vendors' own evaluations and the local-model community report for software-engineering agents, code editing and autocomplete — SWE-bench Verified, Terminal-Bench, DeepSWE — as of the catalogue date. For each VRAM class below we list every model on that list that actually fits at Q4_K_M with an 8K context, so the top row is the pick and the rest are what you give up by choosing it. Numbers are computed, not benchmarked; how the shortlist works.

A coding agent needs context more than most jobs. The "max context" column is the longest window the model can hold on this class at that quantisation — if it reads under 32K, expect to drop a quantisation step or run the model with a q8_0 KV cache before pointing an agent at a whole repository.

Best coding LLM for 8 GB VRAM

7.0 GB usable · speeds for a GeForce RTX 4060

Qwen3.5 9B is the pick: The best coding model for 8–12 GB cards. 4 other models on the coding shortlist also fit, in order of capability:

#ModelWhyParamsQuant · totalMax contextTok/s est.
1 Qwen3.5 9B
Alibaba · Apache 2.0
The best coding model for 8–12 GB cards. 9.65B Q4_K_M · 6.3 GB 31K 30 Tight fit
2 Ornith 1.5 9B
Ornith AI · MIT
A coding-agent reasoning build on the Qwen3.5 9B architecture. 9.41B Q4_K_M · 6.1 GB 35K 31 Tight fit
3 Qwen3.5 4B
Alibaba · Apache 2.0
A coding agent in 3.4 GB — the 6–8 GB answer. 4.66B Q4_K_M · 3.5 GB 120K 63 Runs great
4 Qwen2.5-Coder 7B
Alibaba · Apache 2.0
Still the best fill-in-the-middle autocomplete model under 8B. 7.62B Q4_K_M · 5.3 GB 38K 38 Runs great
5 Qwen3.5 2B
Alibaba · Apache 2.0
Autocomplete on 4 GB. 2.27B Q4_K_M · 2.0 GB 256K 129 Runs great

Best coding LLM for 12 GB VRAM

10.6 GB usable · speeds for a GeForce RTX 3060 12 GB

Qwen3.5 9B is the pick: The best coding model for 8–12 GB cards. 4 other models on the coding shortlist also fit, in order of capability:

#ModelWhyParamsQuant · totalMax contextTok/s est.
1 Qwen3.5 9B
Alibaba · Apache 2.0
The best coding model for 8–12 GB cards. 9.65B Q4_K_M · 6.3 GB 146K 40 Runs great
2 Ornith 1.5 9B
Ornith AI · MIT
A coding-agent reasoning build on the Qwen3.5 9B architecture. 9.41B Q4_K_M · 6.1 GB 150K 41 Runs great
3 Qwen3.5 4B
Alibaba · Apache 2.0
A coding agent in 3.4 GB — the 6–8 GB answer. 4.66B Q4_K_M · 3.5 GB 236K 83 Runs great
4 Qwen2.5-Coder 7B
Alibaba · Apache 2.0
Still the best fill-in-the-middle autocomplete model under 8B. 7.62B Q4_K_M · 5.3 GB 104K 51 Runs great
5 Qwen3.5 2B
Alibaba · Apache 2.0
Autocomplete on 4 GB. 2.27B Q4_K_M · 2.0 GB 256K 171 Runs great

Best coding LLM for 16 GB VRAM

14.4 GB usable · speeds for a GeForce RTX 5080

Qwen3.5 9B is the pick: The best coding model for 8–12 GB cards. 4 other models on the coding shortlist also fit, in order of capability:

#ModelWhyParamsQuant · totalMax contextTok/s est.
1 Qwen3.5 9B
Alibaba · Apache 2.0
The best coding model for 8–12 GB cards. 9.65B Q4_K_M · 6.3 GB 256K 107 Runs great
2 Ornith 1.5 9B
Ornith AI · MIT
A coding-agent reasoning build on the Qwen3.5 9B architecture. 9.41B Q4_K_M · 6.1 GB 256K 110 Runs great
3 Qwen3.5 4B
Alibaba · Apache 2.0
A coding agent in 3.4 GB — the 6–8 GB answer. 4.66B Q4_K_M · 3.5 GB 256K 222 Runs great
4 Qwen2.5-Coder 7B
Alibaba · Apache 2.0
Still the best fill-in-the-middle autocomplete model under 8B. 7.62B Q4_K_M · 5.3 GB 128K 136 Runs great
5 Qwen3.5 2B
Alibaba · Apache 2.0
Autocomplete on 4 GB. 2.27B Q4_K_M · 2.0 GB 256K 455 Runs great

Best coding LLM for 24 GB VRAM

22.4 GB usable · speeds for a GeForce RTX 4090

Qwen3.8 27B is the pick: Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. 11 other models on the coding shortlist also fit, in order of capability:

#ModelWhyParamsQuant · totalMax contextTok/s est.
1 Qwen3.8 27B
Alibaba · Apache 2.0
Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. 27.8B Q4_K_M · 16.7 GB 98K 39 Runs great
2 Qwen3.6 27B
Alibaba · Apache 2.0
77.2 SWE-bench Verified at 27B; the spring-2026 local coding standard. 27.8B Q4_K_M · 16.7 GB 98K 39 Runs great
3 Qwen3.6 35B-A3B
Alibaba · Apache 2.0
Agentic coding at 3B-active speed — the best MoE under 40B for it. 35.9B MoE Q4_K_M · 20.9 GB 82K 126 Tight fit
4 Devstral Small 2 24B
Mistral AI · Apache 2.0
Tuned for software-engineering agents like OpenHands and Cline. 24B Q4_K_M · 15.3 GB 53K 45 Runs great
5 GLM-4.7-Flash 30B-A3B
Z.ai · MIT
MIT-licensed 30B-A3B built for agentic coding; 60–80 tok/s reported on a 4090. 31.2B MoE Q4_K_M · 18.6 GB 82K 139 Runs great
6 Ornith 1.5 35B-A3B
Ornith AI · MIT
Reasoning-first coding build on the Qwen3.6 MoE; thinks before it edits. 35.9B MoE Q4_K_M · 20.9 GB 82K 126 Tight fit
7 Qwen3 Coder 30B-A3B
Alibaba · Apache 2.0
The most-downloaded local code model; 256K window. 30.5B MoE Q4_K_M · 18.5 GB 49K 126 Runs great
8 Qwen3.5 9B
Alibaba · Apache 2.0
The best coding model for 8–12 GB cards. 9.65B Q4_K_M · 6.3 GB 256K 112 Runs great
9 Ornith 1.5 9B
Ornith AI · MIT
A coding-agent reasoning build on the Qwen3.5 9B architecture. 9.41B Q4_K_M · 6.1 GB 256K 115 Runs great
10 Qwen3.5 4B
Alibaba · Apache 2.0
A coding agent in 3.4 GB — the 6–8 GB answer. 4.66B Q4_K_M · 3.5 GB 256K 233 Runs great
11 Qwen2.5-Coder 7B
Alibaba · Apache 2.0
Still the best fill-in-the-middle autocomplete model under 8B. 7.62B Q4_K_M · 5.3 GB 128K 142 Runs great
12 Qwen3.5 2B
Alibaba · Apache 2.0
Autocomplete on 4 GB. 2.27B Q4_K_M · 2.0 GB 256K 478 Runs great

Best coding LLM for 32 GB VRAM

30.4 GB usable · speeds for a GeForce RTX 5090

Qwen3.8 27B is the pick: Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. 11 other models on the coding shortlist also fit, in order of capability:

#ModelWhyParamsQuant · totalMax contextTok/s est.
1 Qwen3.8 27B
Alibaba · Apache 2.0
Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. 27.8B Q4_K_M · 16.7 GB 226K 69 Runs great
2 Qwen3.6 27B
Alibaba · Apache 2.0
77.2 SWE-bench Verified at 27B; the spring-2026 local coding standard. 27.8B Q4_K_M · 16.7 GB 226K 69 Runs great
3 Qwen3.6 35B-A3B
Alibaba · Apache 2.0
Agentic coding at 3B-active speed — the best MoE under 40B for it. 35.9B MoE Q4_K_M · 20.9 GB 256K 225 Runs great
4 Devstral Small 2 24B
Mistral AI · Apache 2.0
Tuned for software-engineering agents like OpenHands and Cline. 24B Q4_K_M · 15.3 GB 104K 80 Runs great
5 GLM-4.7-Flash 30B-A3B
Z.ai · MIT
MIT-licensed 30B-A3B built for agentic coding; 60–80 tok/s reported on a 4090. 31.2B MoE Q4_K_M · 18.6 GB 198K 247 Runs great
6 Ornith 1.5 35B-A3B
Ornith AI · MIT
Reasoning-first coding build on the Qwen3.6 MoE; thinks before it edits. 35.9B MoE Q4_K_M · 20.9 GB 256K 225 Runs great
7 Qwen3 Coder 30B-A3B
Alibaba · Apache 2.0
The most-downloaded local code model; 256K window. 30.5B MoE Q4_K_M · 18.5 GB 134K 225 Runs great
8 Qwen3.5 9B
Alibaba · Apache 2.0
The best coding model for 8–12 GB cards. 9.65B Q4_K_M · 6.3 GB 256K 200 Runs great
9 Ornith 1.5 9B
Ornith AI · MIT
A coding-agent reasoning build on the Qwen3.5 9B architecture. 9.41B Q4_K_M · 6.1 GB 256K 205 Runs great
10 Qwen3.5 4B
Alibaba · Apache 2.0
A coding agent in 3.4 GB — the 6–8 GB answer. 4.66B Q4_K_M · 3.5 GB 256K 414 Runs great
11 Qwen2.5-Coder 7B
Alibaba · Apache 2.0
Still the best fill-in-the-middle autocomplete model under 8B. 7.62B Q4_K_M · 5.3 GB 128K 253 Runs great
12 Qwen3.5 2B
Alibaba · Apache 2.0
Autocomplete on 4 GB. 2.27B Q4_K_M · 2.0 GB 256K 850 Runs great

Best coding LLM for 48 GB VRAM

46.4 GB usable · speeds for a RTX 6000 Ada

Qwen3.8 27B is the pick: Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. 11 other models on the coding shortlist also fit, in order of capability:

#ModelWhyParamsQuant · totalMax contextTok/s est.
1 Qwen3.8 27B
Alibaba · Apache 2.0
Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. 27.8B Q4_K_M · 16.7 GB 256K 37 Runs great
2 Qwen3.6 27B
Alibaba · Apache 2.0
77.2 SWE-bench Verified at 27B; the spring-2026 local coding standard. 27.8B Q4_K_M · 16.7 GB 256K 37 Runs great
3 Qwen3.6 35B-A3B
Alibaba · Apache 2.0
Agentic coding at 3B-active speed — the best MoE under 40B for it. 35.9B MoE Q4_K_M · 20.9 GB 256K 120 Runs great
4 Devstral Small 2 24B
Mistral AI · Apache 2.0
Tuned for software-engineering agents like OpenHands and Cline. 24B Q4_K_M · 15.3 GB 206K 43 Runs great
5 GLM-4.7-Flash 30B-A3B
Z.ai · MIT
MIT-licensed 30B-A3B built for agentic coding; 60–80 tok/s reported on a 4090. 31.2B MoE Q4_K_M · 18.6 GB 198K 133 Runs great
6 Ornith 1.5 35B-A3B
Ornith AI · MIT
Reasoning-first coding build on the Qwen3.6 MoE; thinks before it edits. 35.9B MoE Q4_K_M · 20.9 GB 256K 120 Runs great
7 Qwen3 Coder 30B-A3B
Alibaba · Apache 2.0
The most-downloaded local code model; 256K window. 30.5B MoE Q4_K_M · 18.5 GB 256K 120 Runs great
8 Qwen3.5 9B
Alibaba · Apache 2.0
The best coding model for 8–12 GB cards. 9.65B Q4_K_M · 6.3 GB 256K 107 Runs great
9 Ornith 1.5 9B
Ornith AI · MIT
A coding-agent reasoning build on the Qwen3.5 9B architecture. 9.41B Q4_K_M · 6.1 GB 256K 110 Runs great
10 Qwen3.5 4B
Alibaba · Apache 2.0
A coding agent in 3.4 GB — the 6–8 GB answer. 4.66B Q4_K_M · 3.5 GB 256K 222 Runs great
11 Qwen2.5-Coder 7B
Alibaba · Apache 2.0
Still the best fill-in-the-middle autocomplete model under 8B. 7.62B Q4_K_M · 5.3 GB 128K 136 Runs great
12 Qwen3.5 2B
Alibaba · Apache 2.0
Autocomplete on 4 GB. 2.27B Q4_K_M · 2.0 GB 256K 455 Runs great

Best coding LLM for 64 GB unified memory

48.0 GB usable · speeds for a M4 Max · 64 GB

Qwen3.8 27B is the pick: Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. 11 other models on the coding shortlist also fit, in order of capability:

#ModelWhyParamsQuant · totalMax contextTok/s est.
1 Qwen3.8 27B
Alibaba · Apache 2.0
Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. 27.8B Q4_K_M · 16.7 GB 256K 15 Runs great
2 Qwen3.6 27B
Alibaba · Apache 2.0
77.2 SWE-bench Verified at 27B; the spring-2026 local coding standard. 27.8B Q4_K_M · 16.7 GB 256K 15 Runs great
3 Qwen3.6 35B-A3B
Alibaba · Apache 2.0
Agentic coding at 3B-active speed — the best MoE under 40B for it. 35.9B MoE Q4_K_M · 20.9 GB 256K 58 Runs great
4 Devstral Small 2 24B
Mistral AI · Apache 2.0
Tuned for software-engineering agents like OpenHands and Cline. 24B Q4_K_M · 15.3 GB 216K 17 Runs great
5 GLM-4.7-Flash 30B-A3B
Z.ai · MIT
MIT-licensed 30B-A3B built for agentic coding; 60–80 tok/s reported on a 4090. 31.2B MoE Q4_K_M · 18.6 GB 198K 63 Runs great
6 Ornith 1.5 35B-A3B
Ornith AI · MIT
Reasoning-first coding build on the Qwen3.6 MoE; thinks before it edits. 35.9B MoE Q4_K_M · 20.9 GB 256K 58 Runs great
7 Qwen3 Coder 30B-A3B
Alibaba · Apache 2.0
The most-downloaded local code model; 256K window. 30.5B MoE Q4_K_M · 18.5 GB 256K 58 Runs great
8 Qwen3.5 9B
Alibaba · Apache 2.0
The best coding model for 8–12 GB cards. 9.65B Q4_K_M · 6.3 GB 256K 42 Runs great
9 Ornith 1.5 9B
Ornith AI · MIT
A coding-agent reasoning build on the Qwen3.5 9B architecture. 9.41B Q4_K_M · 6.1 GB 256K 43 Runs great
10 Qwen3.5 4B
Alibaba · Apache 2.0
A coding agent in 3.4 GB — the 6–8 GB answer. 4.66B Q4_K_M · 3.5 GB 256K 87 Runs great
11 Qwen2.5-Coder 7B
Alibaba · Apache 2.0
Still the best fill-in-the-middle autocomplete model under 8B. 7.62B Q4_K_M · 5.3 GB 128K 53 Runs great
12 Qwen3.5 2B
Alibaba · Apache 2.0
Autocomplete on 4 GB. 2.27B Q4_K_M · 2.0 GB 256K 179 Runs great

Best coding LLM for 128 GB unified memory

96.0 GB usable · speeds for a M4 Max · 128 GB

Ling 3.0 Flash 124B-A5B is the pick: SWE-bench Pro 56.6 claimed at 5B active — the 128 GB-class coding pick if you cannot fit DeepSeek V4. 12 other models on the coding shortlist also fit, in order of capability:

#ModelWhyParamsQuant · totalMax contextTok/s est.
1 Ling 3.0 Flash 124B-A5B
inclusionAI · MIT
SWE-bench Pro 56.6 claimed at 5B active — the 128 GB-class coding pick if you cannot fit DeepSeek V4. 124B MoE Q4_K_M · 70.4 GB 256K 50 Runs great
2 Qwen3.8 27B
Alibaba · Apache 2.0
Terminal-Bench 73, DeepSWE 42 — a generation ahead of anything else that fits 24 GB. 27.8B Q4_K_M · 16.7 GB 256K 20 Runs great
3 Qwen3.6 27B
Alibaba · Apache 2.0
77.2 SWE-bench Verified at 27B; the spring-2026 local coding standard. 27.8B Q4_K_M · 16.7 GB 256K 20 Runs great
4 Qwen3.6 35B-A3B
Alibaba · Apache 2.0
Agentic coding at 3B-active speed — the best MoE under 40B for it. 35.9B MoE Q4_K_M · 20.9 GB 256K 77 Runs great
5 Devstral Small 2 24B
Mistral AI · Apache 2.0
Tuned for software-engineering agents like OpenHands and Cline. 24B Q4_K_M · 15.3 GB 384K 23 Runs great
6 GLM-4.7-Flash 30B-A3B
Z.ai · MIT
MIT-licensed 30B-A3B built for agentic coding; 60–80 tok/s reported on a 4090. 31.2B MoE Q4_K_M · 18.6 GB 198K 84 Runs great
7 Ornith 1.5 35B-A3B
Ornith AI · MIT
Reasoning-first coding build on the Qwen3.6 MoE; thinks before it edits. 35.9B MoE Q4_K_M · 20.9 GB 256K 77 Runs great
8 Qwen3 Coder 30B-A3B
Alibaba · Apache 2.0
The most-downloaded local code model; 256K window. 30.5B MoE Q4_K_M · 18.5 GB 256K 77 Runs great
9 Qwen3.5 9B
Alibaba · Apache 2.0
The best coding model for 8–12 GB cards. 9.65B Q4_K_M · 6.3 GB 256K 56 Runs great
10 Ornith 1.5 9B
Ornith AI · MIT
A coding-agent reasoning build on the Qwen3.5 9B architecture. 9.41B Q4_K_M · 6.1 GB 256K 58 Runs great
11 Qwen3.5 4B
Alibaba · Apache 2.0
A coding agent in 3.4 GB — the 6–8 GB answer. 4.66B Q4_K_M · 3.5 GB 256K 116 Runs great
12 Qwen2.5-Coder 7B
Alibaba · Apache 2.0
Still the best fill-in-the-middle autocomplete model under 8B. 7.62B Q4_K_M · 5.3 GB 128K 71 Runs great
13 Qwen3.5 2B
Alibaba · Apache 2.0
Autocomplete on 4 GB. 2.27B Q4_K_M · 2.0 GB 256K 239 Runs great

The whole coding shortlist, and what each one needs

Every model on the list, best first, with the smallest discrete card that runs it with 15% headroom at the recommended quantisation and 8K context.

#ModelParamsWeightsNeedsOllama tag
1 DeepSeek V4.1 Flash 552B-A16B New
Twice V4 Flash's backbone (552B, 16B active per generated token) plus 196B of Engram lookup tables (189 GiB at FP8) that DwarfStar streams from SSD, so they are not counted here. A 512 GB Mac model, and as of October 2026 no mainline llama.cpp or Ollama build: DwarfStar on a Mac, vLLM across four GPUs. DeepSeek puts the cache at 890 bytes/token; modelled conservatively as 4 latent layers plus a 128-token window.
552B MoE 310.4 GB @ Q4_K_M more than one card —
2 Kimi K2.6 1T-A32B
The open coding-agent benchmark leader of spring 2026 (80.2 SWE-bench). A 512 GB Mac Studio pair, or a ceiling.
1027B MoE 577.5 GB @ Q4_K_M more than one card —
3 GLM-5.3 744B-A40B
Same base as GLM-5.2, new post-training: Z.ai's most capable open-weights coder (Terminal-Bench 3.0 28.3, DeepSWE 66.9). Not MIT any more — its own GLM-5.3 licence. Same ceiling as 5.2: 512 GB of unified memory at Q4.
753B MoE 423.4 GB @ Q4_K_M more than one card —
4 GLM-5.3-Flash 320B-A18B
The first natively multimodal GLM-5 and the first hybrid: 34 linear-attention blocks and 11 sparse-attention blocks with a 512-wide latent cache, so a 1M window stays affordable. Z.ai says it beats GLM-5.2 at 18B active; MIT.
321B MoE 180.5 GB @ Q4_K_M more than one card —
5 DeepSeek V4 Flash 284B-A13B
The V4 for 192–256 GB machines: Q3 squeezes into 192 GB, Q4 wants 256 GB, and 128 GB falls short even at Q2. Cache is modelled as a 576-wide latent; V4 compresses it further at long context, so this is conservative.
284B MoE 159.7 GB @ Q4_K_M 192 GB card —
6 Qwen3.8-Flash-Next 180B-A6B
The open preview of the Qwen4 architecture: a 125B-A6B hybrid (Gated DeltaNet + sparse attention, KV cache on 12 of 48 blocks) plus a 51B n-gram embedding and a 4B draft head — 180B on disk, 6B active. The hosted "Qwen3.8-Flash" is this model with a 1M window.
180B MoE 101.2 GB @ Q4_K_M 141 GB card qwen3.8-flash-next:125b-a6b-nvfp4
7 Ling 3.0 Flash 124B-A5B
A 124B hybrid (5 linear-attention layers per MLA layer) with 5.1B active: SWE-bench Pro 56.6 and AIME 93 claimed. Built for 96–128 GB machines.
124B MoE 69.7 GB @ Q4_K_M 141 GB card —
8 Qwen3.8 27B
The current default local Qwen: dense 27B, text + image + video, 262K context. Only 16 of its 64 blocks keep a KV cache, so long context is cheap.
27.8B 15.6 GB @ Q4_K_M 24 GB card qwen3.8:27b
9 Qwen3.6 27B
The 24 GB coding pick of spring 2026 (77.2 SWE-bench Verified). Same shape as 3.8, one generation behind.
27.8B 15.6 GB @ Q4_K_M 24 GB card qwen3.6:27b
10 Qwen3.6 35B-A3B
Mixture of experts with ~3B active: the fastest serious model a 24 GB card runs, and the best MoE under 40B on agentic coding.
35.9B MoE 20.2 GB @ Q4_K_M 32 GB card qwen3.6:35b-a3b
11 Devstral Small 2 24B
Built for software-engineering agents (OpenHands, Cline). Dense 24B, 384K window. Not in the Ollama library.
24B 13.5 GB @ Q4_K_M 20 GB card —
12 GLM-4.7-Flash 30B-A3B
MIT-licensed 30B-A3B tuned for agentic coding, with a DeepSeek-style latent KV cache. 60–80 tok/s reported on a 4090.
31.2B MoE 17.5 GB @ Q4_K_M 24 GB card glm-4.7-flash:latest
13 Ornith 1.5 35B-A3B
A reasoning-first MIT build on the Qwen3.6 35B-A3B architecture (thinks before every answer). Same VRAM as its base.
35.9B MoE 20.2 GB @ Q4_K_M 32 GB card ornith-1.5:35b
14 Qwen3 Coder 30B-A3B
Agentic coding MoE with a 256K native window. Still the most-downloaded local code model.
30.5B MoE 17.1 GB @ Q4_K_M 24 GB card qwen3-coder:30b
15 Qwen3.5 9B
The default for 8–12 GB cards in 2026: beats every older 8B on every published benchmark, with vision.
9.65B 5.4 GB @ Q4_K_M 10 GB card qwen3.5:9b
16 Ornith 1.5 9B
The small Ornith: a coding-agent reasoning build on the Qwen3.5 9B architecture. Same VRAM as its base, thinks before every answer.
9.41B 5.3 GB @ Q4_K_M 10 GB card ornith-1.5:9b
17 Qwen3.5 4B
The 8 GB coding agent. Q4 lands near 3.4 GB, leaving room for a real context window.
4.66B 2.6 GB @ Q4_K_M 8 GB card qwen3.5:4b
18 Qwen2.5-Coder 7B
The standard local autocomplete model — small enough to keep resident all day, and still the best FIM model under 8B.
7.62B 4.3 GB @ Q4_K_M 8 GB card qwen2.5-coder:7b
19 Qwen3.5 2B
Phone-class, and multimodal. Replaces Llama 3.2 3B as the "it runs on anything" answer.
2.27B 1.3 GB @ Q4_K_M 8 GB card qwen3.5:2b

Autocomplete and agents want different things. Fill-in-the-middle completion is latency-bound and happy with a 4–9B model; an agent that edits files and runs tests wants the strongest model you can hold with a 32K+ window, and will wait for it. If you are unsure which card to buy for this, the GPU ranking says what each class unlocks.