Can your hardware run it?
Find the open-weight models your computer can actually run — and which one to pick for coding, reasoning, vision or chat. Choose your GPU or Apple Silicon chip: we size the weights at your quantisation, add the KV cache at your context length, estimate tokens per second, and recommend a model per job.
Any org/model id or hf.co link — not just the 79 we have verified. We read its config.json, work out the parameters, layers, KV heads, context window and experts, and size it against your GPU. Ollama tags and model names work too.
New this month
16 models released in the 31 days before 28 Aug 2026 · all models, newest firstThe first natively multimodal GLM-5 and the first hybrid: 34 linear-attention blocks and 11 sparse-attention blocks with a 512-wide latent cache, so a 1M window stays affordable. Z.ai says it beats GLM-5.2 at 18B active; MIT.
Same base as GLM-5.2, new post-training: Z.ai's most capable open-weights coder (Terminal-Bench 3.0 28.3, DeepSWE 66.9). Not MIT any more — its own GLM-5.3 licence. Same ceiling as 5.2: 512 GB of unified memory at Q4.
The open preview of the Qwen4 architecture: a 125B-A6B hybrid (Gated DeltaNet + sparse attention, KV cache on 12 of 48 blocks) plus a 51B n-gram embedding and a 4B draft head — 180B on disk, 6B active. The hosted "Qwen3.8-Flash" is this model with a 1M window.
The small Ornith: a coding-agent reasoning build on the Qwen3.5 9B architecture. Same VRAM as its base, thinks before every answer.
A reasoning-first MIT build on the Qwen3.6 35B-A3B architecture (thinks before every answer). Same VRAM as its base.
The flagship Ornith on the Qwen3.5 397B-A17B architecture, MIT-licensed. A 256 GB Mac Studio at Q4, and it is in the Ollama library.
The current default local Qwen: dense 27B, text + image + video, 262K context. Only 16 of its 64 blocks keep a KV cache, so long context is cheap.
Japan’s national-institute reasoning model, Japanese and English. A plain dense Llama-style 33B: Q4 is a tight 24 GB fit.
The 0813 refresh of the V4 flagship. Included as the honest ceiling; a terabyte of weights at Q4.
The first open Qwen-Max-class flagship. Listed as the honest ceiling: nothing short of a rack runs it.
Mamba-2 + MoE hybrid built for the execution layer of agents: only 6 attention blocks, so the KV cache is almost free. Weights, data and recipe all open.
An 8B MoE with 1.3B active and a latent KV cache on only 6 of 24 layers — reasoning and tool use sized for Apple Silicon and edge boxes.
Meta's first open weights since Llama 4: a dense 30B distilled from Muse Spark for always-on local agents. Two KV heads keep the cache small.
A 124B hybrid (5 linear-attention layers per MLA layer) with 5.1B active: SWE-bench Pro 56.6 and AIME 93 claimed. Built for 96–128 GB machines.
The V4 that 128 GB machines can actually run at Q3. Cache is modelled as a 576-wide latent; V4 compresses it further at long context, so this is conservative.
Convolution-heavy hybrid for CPUs and NPUs: 22 of 30 blocks keep no KV cache at all.