Models

76 open-weight models, sized

Updated 21 Aug 2026 5 new this week, 14 this month

Parameter counts, layer counts and attention shapes are read from each model's own config, because the KV-cache arithmetic depends on them exactly. Sizes below are at Q4_K_M and 8K context against a GeForce RTX 3060 12 GB. Missing a model that shipped this week? Open an issue — the table is hand-maintained.

ModelReleasedParamsQuantWeightsMax contextLicenceOn this rig
Qwen3 0.6B
Useful mostly as a speculative-decoding draft model for its larger siblings.
Apr 2025 0.6B Q4_K_M 0.3 GB 32K Apache 2.0 646 tok/s Check · Hardware
Qwen3.5 0.8B
Draft model for speculative decoding, or a classifier that fits in 1 GB.
28 Feb 2026 0.87B Q4_K_M 0.5 GB 256K Apache 2.0 445 tok/s Check · Hardware
Gemma 3 1B
Text-only. Single KV head makes its cache almost free at long context.
Mar 2025 1.0B Q4_K_M 0.6 GB 32K Gemma Terms of Use 388 tok/s Check · Hardware
Llama 3.2 1B Instruct
The smallest Llama worth running. Fits anywhere, including phones and 4 GB cards.
Sep 2024 1.24B Q4_K_M 0.7 GB 128K Llama 3.2 Community 313 tok/s Check · Hardware
Qwen3 1.7B
Punches above its size on structured tasks, with optional thinking mode.
Apr 2025 1.72B Q4_K_M 1.0 GB 32K Apache 2.0 225 tok/s Check · Hardware
Qwen3.5 2B
Phone-class, and multimodal. Replaces Llama 3.2 3B as the "it runs on anything" answer.
28 Feb 2026 2.27B Q4_K_M 1.3 GB 256K Apache 2.0 171 tok/s Check · Hardware
LFM2.5 2.6B New
Convolution-heavy hybrid for CPUs and NPUs: 22 of 30 blocks keep no KV cache at all.
28 Jul 2026 2.7B Q4_K_M 1.5 GB 128K LFM Open License v1.0 144 tok/s Check · Hardware
Llama 3.2 3B Instruct
The 2024 "it just runs" model for 8 GB laptops. Qwen3.5 4B does the same job better now.
Sep 2024 3.21B Q4_K_M 1.8 GB 128K Llama 3.2 Community 121 tok/s Check · Hardware
Granite 4.1 3B
Dense, small, enterprise-flavoured: tool calling and instruction following, no thinking mode.
29 Apr 2026 3.4B Q4_K_M 1.9 GB 128K Apache 2.0 114 tok/s Check · Hardware
Phi-4-mini 3.8B
MIT-licensed, dense, and unusually strong on instruction following for its size.
Feb 2025 3.84B Q4_K_M 2.2 GB 128K MIT 101 tok/s Check · Hardware
Ministral 3 3B
Edge model with a vision encoder and a 256K window. Apache 2.0.
Dec 2025 3.85B Q4_K_M 2.2 GB 256K Apache 2.0 101 tok/s Check · Hardware
Qwen3 4B
The 2025 sweet spot for 8 GB cards with reasoning traces. Qwen3.5 4B adds vision and 8× the context.
Apr 2025 4.02B Q4_K_M 2.3 GB 32K Apache 2.0 96 tok/s Check · Hardware
Gemma 3 4B
Vision-capable at 4B. Superseded by Gemma 4 E4B, still everywhere.
Mar 2025 4.3B Q4_K_M 2.4 GB 128K Gemma Terms of Use 90 tok/s Check · Hardware
Qwen3.5 4B
The 8 GB coding agent. Q4 lands near 3.4 GB, leaving room for a real context window.
28 Feb 2026 4.66B Q4_K_M 2.6 GB 256K Apache 2.0 83 tok/s Check · Hardware
Gemma 4 E2B
"E2B" is 2.3B effective, but the file holds 5B because of per-layer embeddings — size it as 5B. Text, image and audio in.
2 Apr 2026 5.1B Q4_K_M 2.9 GB 128K Apache 2.0 76 tok/s Check · Hardware
Qwen2.5-Coder 7B
The standard local autocomplete model — small enough to keep resident all day, and still the best FIM model under 8B.
Nov 2024 7.62B Q4_K_M 4.3 GB 128K Apache 2.0 51 tok/s Check · Hardware
Ling 3.0 Tiny 7.9B-A1.3B New
An 8B MoE with 1.3B active and a latent KV cache on only 6 of 24 layers — reasoning and tool use sized for Apple Silicon and edge boxes.
10 Aug 2026 7.9B MoE Q4_K_M 4.4 GB 128K MIT 115 tok/s Check · Hardware
Gemma 4 E4B
The laptop Gemma. 4.5B effective, 8B on disk; a single KV head per window layer keeps its cache tiny.
2 Apr 2026 8.0B Q4_K_M 4.5 GB 128K Apache 2.0 48 tok/s Check · Hardware
LFM2.5 8B-A1B
An 8B MoE with ~1.5B active, aimed at laptops without a GPU. Licence is permissive below $10M revenue.
28 May 2026 8.47B MoE Q4_K_M 4.8 GB 125K LFM Open License v1.0 99 tok/s Check · Hardware
Granite 4.1 8B
Matches the old Granite 4.0 32B MoE at a quarter of the size. Fast, Apache 2.0, no reasoning traces.
29 Apr 2026 8.79B Q4_K_M 4.9 GB 128K Apache 2.0 44 tok/s Check · Hardware
Ministral 3 8B
Mistral's 8B with images in. Plain GQA, so budget more KV cache than Qwen3.5 9B at the same context.
Dec 2025 8.92B Q4_K_M 5.0 GB 256K Apache 2.0 43 tok/s Check · Hardware
Ornith 1.5 9B New this week
The small Ornith: a coding-agent reasoning build on the Qwen3.5 9B architecture. Same VRAM as its base, thinks before every answer.
19 Aug 2026 9.41B Q4_K_M 5.3 GB 256K MIT 41 tok/s Check · Hardware
Qwen3.5 9B
The default for 8–12 GB cards in 2026: beats every older 8B on every published benchmark, with vision.
28 Feb 2026 9.65B Q4_K_M 5.4 GB 256K Apache 2.0 40 tok/s Check · Hardware
gpt-oss 20B
Ships natively in MXFP4, so the 4-bit weights are the reference weights, not a lossy copy. Fits 16 GB.
Aug 2025 20.9B MoE MXFP4 10.8 GB 128K Apache 2.0 offload Check · Hardware
Gemma 4 26B-A4B
Mixture of experts with 3.8B active. Slower to think than Qwen3.6 35B-A3B, faster to answer, and it sees images.
2 Apr 2026 26.5B MoE Q4_K_M 14.9 GB 256K Apache 2.0 offload Check · Hardware
Qwen3 30B-A3B
The MoE that made "3B active" a category. Qwen3.6 35B-A3B is its direct replacement.
Apr 2025 30.5B MoE Q4_K_M 17.1 GB 128K Apache 2.0 offload Check · Hardware
Qwen3 Coder 30B-A3B
Agentic coding MoE with a 256K native window. Still the most-downloaded local code model.
Jul 2025 30.5B MoE Q4_K_M 17.1 GB 256K Apache 2.0 offload Check · Hardware
GLM-4.7-Flash 30B-A3B
MIT-licensed 30B-A3B tuned for agentic coding, with a DeepSeek-style latent KV cache. 60–80 tok/s reported on a 4090.
20 Jan 2026 31.2B MoE Q4_K_M 17.5 GB 198K MIT offload Check · Hardware
Nemotron 3.5 Lightning 30B-A3B New
Mamba-2 + MoE hybrid built for the execution layer of agents: only 6 attention blocks, so the KV cache is almost free. Weights, data and recipe all open.
11 Aug 2026 31.6B MoE Q4_K_M 17.8 GB 256K OpenMDW-1.1 offload Check · Hardware
Qwen3.6 35B-A3B
Mixture of experts with ~3B active: the fastest serious model a 24 GB card runs, and the best MoE under 40B on agentic coding.
16 Apr 2026 35.9B MoE Q4_K_M 20.2 GB 256K Apache 2.0 offload Check · Hardware
Ornith 1.5 35B-A3B New this week
A reasoning-first MIT build on the Qwen3.6 35B-A3B architecture (thinks before every answer). Same VRAM as its base.
19 Aug 2026 35.9B MoE Q4_K_M 20.2 GB 256K MIT offload Check · Hardware

Weight sizes are computed from the parameter count and the quantisation's effective bits per weight, not read off a file listing — expect them to land within a few percent of the GGUF you actually download. The method, in full.