Models
76 open-weight models, sized
Updated 21 Aug 2026 5 new this week, 14 this month
Parameter counts, layer counts and attention shapes are read from each model's own config, because the KV-cache arithmetic depends on them exactly. Sizes below are at Q4_K_M and 8K context against a GeForce RTX 3060 12 GB. Missing a model that shipped this week? Open an issue — the table is hand-maintained.
| Model | Released | Params | Quant | Weights | Max context | Licence | On this rig | |
|---|---|---|---|---|---|---|---|---|
| Qwen3 0.6B Useful mostly as a speculative-decoding draft model for its larger siblings. |
Apr 2025 | 0.6B | Q4_K_M | 0.3 GB | 32K | Apache 2.0 | 646 tok/s | Check · Hardware |
| Qwen3.5 0.8B Draft model for speculative decoding, or a classifier that fits in 1 GB. |
28 Feb 2026 | 0.87B | Q4_K_M | 0.5 GB | 256K | Apache 2.0 | 445 tok/s | Check · Hardware |
| Gemma 3 1B Text-only. Single KV head makes its cache almost free at long context. |
Mar 2025 | 1.0B | Q4_K_M | 0.6 GB | 32K | Gemma Terms of Use | 388 tok/s | Check · Hardware |
| Llama 3.2 1B Instruct The smallest Llama worth running. Fits anywhere, including phones and 4 GB cards. |
Sep 2024 | 1.24B | Q4_K_M | 0.7 GB | 128K | Llama 3.2 Community | 313 tok/s | Check · Hardware |
| Qwen3 1.7B Punches above its size on structured tasks, with optional thinking mode. |
Apr 2025 | 1.72B | Q4_K_M | 1.0 GB | 32K | Apache 2.0 | 225 tok/s | Check · Hardware |
| Qwen3.5 2B Phone-class, and multimodal. Replaces Llama 3.2 3B as the "it runs on anything" answer. |
28 Feb 2026 | 2.27B | Q4_K_M | 1.3 GB | 256K | Apache 2.0 | 171 tok/s | Check · Hardware |
| LFM2.5 2.6B New Convolution-heavy hybrid for CPUs and NPUs: 22 of 30 blocks keep no KV cache at all. |
28 Jul 2026 | 2.7B | Q4_K_M | 1.5 GB | 128K | LFM Open License v1.0 | 144 tok/s | Check · Hardware |
| Llama 3.2 3B Instruct The 2024 "it just runs" model for 8 GB laptops. Qwen3.5 4B does the same job better now. |
Sep 2024 | 3.21B | Q4_K_M | 1.8 GB | 128K | Llama 3.2 Community | 121 tok/s | Check · Hardware |
| Granite 4.1 3B Dense, small, enterprise-flavoured: tool calling and instruction following, no thinking mode. |
29 Apr 2026 | 3.4B | Q4_K_M | 1.9 GB | 128K | Apache 2.0 | 114 tok/s | Check · Hardware |
| Phi-4-mini 3.8B MIT-licensed, dense, and unusually strong on instruction following for its size. |
Feb 2025 | 3.84B | Q4_K_M | 2.2 GB | 128K | MIT | 101 tok/s | Check · Hardware |
| Ministral 3 3B Edge model with a vision encoder and a 256K window. Apache 2.0. |
Dec 2025 | 3.85B | Q4_K_M | 2.2 GB | 256K | Apache 2.0 | 101 tok/s | Check · Hardware |
| Qwen3 4B The 2025 sweet spot for 8 GB cards with reasoning traces. Qwen3.5 4B adds vision and 8× the context. |
Apr 2025 | 4.02B | Q4_K_M | 2.3 GB | 32K | Apache 2.0 | 96 tok/s | Check · Hardware |
| Gemma 3 4B Vision-capable at 4B. Superseded by Gemma 4 E4B, still everywhere. |
Mar 2025 | 4.3B | Q4_K_M | 2.4 GB | 128K | Gemma Terms of Use | 90 tok/s | Check · Hardware |
| Qwen3.5 4B The 8 GB coding agent. Q4 lands near 3.4 GB, leaving room for a real context window. |
28 Feb 2026 | 4.66B | Q4_K_M | 2.6 GB | 256K | Apache 2.0 | 83 tok/s | Check · Hardware |
| Gemma 4 E2B "E2B" is 2.3B effective, but the file holds 5B because of per-layer embeddings — size it as 5B. Text, image and audio in. |
2 Apr 2026 | 5.1B | Q4_K_M | 2.9 GB | 128K | Apache 2.0 | 76 tok/s | Check · Hardware |
| Qwen2.5-Coder 7B The standard local autocomplete model — small enough to keep resident all day, and still the best FIM model under 8B. |
Nov 2024 | 7.62B | Q4_K_M | 4.3 GB | 128K | Apache 2.0 | 51 tok/s | Check · Hardware |
| Ling 3.0 Tiny 7.9B-A1.3B New An 8B MoE with 1.3B active and a latent KV cache on only 6 of 24 layers — reasoning and tool use sized for Apple Silicon and edge boxes. |
10 Aug 2026 | 7.9B MoE | Q4_K_M | 4.4 GB | 128K | MIT | 115 tok/s | Check · Hardware |
| Gemma 4 E4B The laptop Gemma. 4.5B effective, 8B on disk; a single KV head per window layer keeps its cache tiny. |
2 Apr 2026 | 8.0B | Q4_K_M | 4.5 GB | 128K | Apache 2.0 | 48 tok/s | Check · Hardware |
| LFM2.5 8B-A1B An 8B MoE with ~1.5B active, aimed at laptops without a GPU. Licence is permissive below $10M revenue. |
28 May 2026 | 8.47B MoE | Q4_K_M | 4.8 GB | 125K | LFM Open License v1.0 | 99 tok/s | Check · Hardware |
| Granite 4.1 8B Matches the old Granite 4.0 32B MoE at a quarter of the size. Fast, Apache 2.0, no reasoning traces. |
29 Apr 2026 | 8.79B | Q4_K_M | 4.9 GB | 128K | Apache 2.0 | 44 tok/s | Check · Hardware |
| Ministral 3 8B Mistral's 8B with images in. Plain GQA, so budget more KV cache than Qwen3.5 9B at the same context. |
Dec 2025 | 8.92B | Q4_K_M | 5.0 GB | 256K | Apache 2.0 | 43 tok/s | Check · Hardware |
| Ornith 1.5 9B New this week The small Ornith: a coding-agent reasoning build on the Qwen3.5 9B architecture. Same VRAM as its base, thinks before every answer. |
19 Aug 2026 | 9.41B | Q4_K_M | 5.3 GB | 256K | MIT | 41 tok/s | Check · Hardware |
| Qwen3.5 9B The default for 8–12 GB cards in 2026: beats every older 8B on every published benchmark, with vision. |
28 Feb 2026 | 9.65B | Q4_K_M | 5.4 GB | 256K | Apache 2.0 | 40 tok/s | Check · Hardware |
| gpt-oss 20B Ships natively in MXFP4, so the 4-bit weights are the reference weights, not a lossy copy. Fits 16 GB. |
Aug 2025 | 20.9B MoE | MXFP4 | 10.8 GB | 128K | Apache 2.0 | offload | Check · Hardware |
| Gemma 4 26B-A4B Mixture of experts with 3.8B active. Slower to think than Qwen3.6 35B-A3B, faster to answer, and it sees images. |
2 Apr 2026 | 26.5B MoE | Q4_K_M | 14.9 GB | 256K | Apache 2.0 | offload | Check · Hardware |
| Qwen3 30B-A3B The MoE that made "3B active" a category. Qwen3.6 35B-A3B is its direct replacement. |
Apr 2025 | 30.5B MoE | Q4_K_M | 17.1 GB | 128K | Apache 2.0 | offload | Check · Hardware |
| Qwen3 Coder 30B-A3B Agentic coding MoE with a 256K native window. Still the most-downloaded local code model. |
Jul 2025 | 30.5B MoE | Q4_K_M | 17.1 GB | 256K | Apache 2.0 | offload | Check · Hardware |
| GLM-4.7-Flash 30B-A3B MIT-licensed 30B-A3B tuned for agentic coding, with a DeepSeek-style latent KV cache. 60–80 tok/s reported on a 4090. |
20 Jan 2026 | 31.2B MoE | Q4_K_M | 17.5 GB | 198K | MIT | offload | Check · Hardware |
| Nemotron 3.5 Lightning 30B-A3B New Mamba-2 + MoE hybrid built for the execution layer of agents: only 6 attention blocks, so the KV cache is almost free. Weights, data and recipe all open. |
11 Aug 2026 | 31.6B MoE | Q4_K_M | 17.8 GB | 256K | OpenMDW-1.1 | offload | Check · Hardware |
| Qwen3.6 35B-A3B Mixture of experts with ~3B active: the fastest serious model a 24 GB card runs, and the best MoE under 40B on agentic coding. |
16 Apr 2026 | 35.9B MoE | Q4_K_M | 20.2 GB | 256K | Apache 2.0 | offload | Check · Hardware |
| Ornith 1.5 35B-A3B New this week A reasoning-first MIT build on the Qwen3.6 35B-A3B architecture (thinks before every answer). Same VRAM as its base. |
19 Aug 2026 | 35.9B MoE | Q4_K_M | 20.2 GB | 256K | MIT | offload | Check · Hardware |
Weight sizes are computed from the parameter count and the quantisation's effective bits per weight, not read off a file listing — expect them to land within a few percent of the GGUF you actually download. The method, in full.