Models

76 open-weight models, sized

Updated 21 Aug 2026 5 new this week, 14 this month

Parameter counts, layer counts and attention shapes are read from each model's own config, because the KV-cache arithmetic depends on them exactly. Sizes below are at Q4_K_M and 8K context against a Radeon RX 7900 XTX. Missing a model that shipped this week? Open an issue — the table is hand-maintained.

ModelReleasedParamsQuantWeightsMax contextLicenceOn this rig
Olmo 3 7B Instruct
Fully open — training data and code included. Full multi-head attention, so its cache is 4× a GQA 7B at the same context.
20 Nov 2025 7.3B Q4_K_M 4.1 GB 64K Apache 2.0 142 tok/s Check · Hardware
Llama 3.1 8B Instruct
Still the most widely deployed local model, with the largest fine-tune ecosystem. Not the strongest 8B any more.
Jul 2024 8.03B Q4_K_M 4.5 GB 128K Llama 3.1 Community 129 tok/s Check · Hardware
Ministral 3 8B
Mistral's 8B with images in. Plain GQA, so budget more KV cache than Qwen3.5 9B at the same context.
Dec 2025 8.92B Q4_K_M 5.0 GB 256K Apache 2.0 116 tok/s Check · Hardware
Gemma 4 12B
The "unified" Gemma 4: text, image and audio in one 12B that fits a 12 GB card at Q4. 140+ languages.
29 May 2026 12B Q4_K_M 6.7 GB 256K Apache 2.0 86 tok/s Check · Hardware
Mistral NeMo 12B
Multilingual 12B with a 128K window, built with NVIDIA. A roleplay and fiction favourite that refuses to die.
Jul 2024 12.2B Q4_K_M 6.9 GB 128K Apache 2.0 85 tok/s Check · Hardware
Ministral 3 14B
The largest Ministral. A 12 GB card runs it at Q4 with a few gigabytes to spare.
Dec 2025 13.9B Q4_K_M 7.8 GB 256K Apache 2.0 74 tok/s Check · Hardware
Mistral Small 3.2 24B
Apache-2.0, vision-capable, and the most 24 GB-friendly of the 2025 generalists.
Jun 2025 23.6B Q4_K_M 13.3 GB 128K Apache 2.0 44 tok/s Check · Hardware
Gemma 4 26B-A4B
Mixture of experts with 3.8B active. Slower to think than Qwen3.6 35B-A3B, faster to answer, and it sees images.
2 Apr 2026 26.5B MoE Q4_K_M 14.9 GB 256K Apache 2.0 105 tok/s Check · Hardware
Gemma 3 27B
The 2025 single-GPU generalist with vision. Its Gemma-licence terms are the reason to prefer Gemma 4 now.
Mar 2025 27.4B Q4_K_M 15.4 GB 128K Gemma Terms of Use 38 tok/s Check · Hardware
Gemma 4 31B
The dense flagship: strongest maths of the 24–32 GB class (89% AIME), clean prose, vision. Q4 is a tight 24 GB fit.
2 Apr 2026 31.3B Q4_K_M 17.6 GB 256K Apache 2.0 33 tok/s Check · Hardware
Olmo 3.1 32B Instruct
The largest fully open model you can audit end to end. Q4 fits 24 GB, tightly.
10 Dec 2025 32.2B Q4_K_M 18.1 GB 64K Apache 2.0 32 tok/s Check · Hardware
Qwen3 32B
The classic 24 GB target, and still the strongest local translator under 70B. Qwen3.8 27B is smaller and better at everything else.
Apr 2025 32.8B Q4_K_M 18.4 GB 128K Apache 2.0 32 tok/s Check · Hardware
Llama 3.3 70B Instruct
Still the creative-writing favourite: consistent voice, takes direction. Needs 48 GB to sit comfortably on GPU at Q4.
Dec 2024 70.6B Q4_K_M 39.7 GB 128K Llama 3.3 Community offload Check · Hardware

Weight sizes are computed from the parameter count and the quantisation's effective bits per weight, not read off a file listing — expect them to land within a few percent of the GGUF you actually download. The method, in full.