Updated 28 Aug 2026 3 new this week, 16 this month

Can your hardware run it?

Find the open-weight models your computer can actually run — and which one to pick for coding, reasoning, vision or chat. Choose your GPU or Apple Silicon chip: we size the weights at your quantisation, add the KV cache at your context length, estimate tokens per second, and recommend a model per job.

Paste any Hugging Face model

Any org/model id or hf.co link — not just the 79 we have verified. We read its config.json, work out the parameters, layers, KV heads, context window and experts, and size it against your GPU. Ollama tags and model names work too.


New this month

16 models released in the 31 days before 28 Aug 2026 · all models, newest first
GLM-5.3-Flash 320B-A18BNew this week

The first natively multimodal GLM-5 and the first hybrid: 34 linear-attention blocks and 11 sparse-attention blocks with a 512-wide latent cache, so a 1M window stays affordable. Z.ai says it beats GLM-5.2 at 18B active; MIT.

Z.ai · 25 Aug 2026 · MIT
GLM-5.3 744B-A40BNew this week

Same base as GLM-5.2, new post-training: Z.ai's most capable open-weights coder (Terminal-Bench 3.0 28.3, DeepSWE 66.9). Not MIT any more — its own GLM-5.3 licence. Same ceiling as 5.2: 512 GB of unified memory at Q4.

Z.ai · 25 Aug 2026 · GLM-5.3 License
Qwen3.8-Flash-Next 180B-A6BNew this week

The open preview of the Qwen4 architecture: a 125B-A6B hybrid (Gated DeltaNet + sparse attention, KV cache on 12 of 48 blocks) plus a 51B n-gram embedding and a 4B draft head — 180B on disk, 6B active. The hosted "Qwen3.8-Flash" is this model with a 1M window.

Alibaba · 24 Aug 2026 · Qwen Community 1.0
Ornith 1.5 9BNew

The small Ornith: a coding-agent reasoning build on the Qwen3.5 9B architecture. Same VRAM as its base, thinks before every answer.

Ornith AI · 19 Aug 2026 · MIT
Ornith 1.5 35B-A3BNew

A reasoning-first MIT build on the Qwen3.6 35B-A3B architecture (thinks before every answer). Same VRAM as its base.

Ornith AI · 19 Aug 2026 · MIT
Ornith 1.5 397B-A17BNew

The flagship Ornith on the Qwen3.5 397B-A17B architecture, MIT-licensed. A 256 GB Mac Studio at Q4, and it is in the Ollama library.

Ornith AI · 19 Aug 2026 · MIT
Qwen3.8 27BNew

The current default local Qwen: dense 27B, text + image + video, 262K context. Only 16 of its 64 blocks keep a KV cache, so long context is cheap.

Alibaba · 14 Aug 2026 · Apache 2.0
LLM-jp 4 33B ThinkingNew

Japan’s national-institute reasoning model, Japanese and English. A plain dense Llama-style 33B: Q4 is a tight 24 GB fit.

NII (Japan) · 14 Aug 2026 · Apache 2.0
DeepSeek V4 Pro 1.6T-A49BNew

The 0813 refresh of the V4 flagship. Included as the honest ceiling; a terabyte of weights at Q4.

DeepSeek · 13 Aug 2026 · MIT
Qwen3.8 2.4T-A95BNew

The first open Qwen-Max-class flagship. Listed as the honest ceiling: nothing short of a rack runs it.

Alibaba · 12 Aug 2026 · Qwen3.8-Max License
Nemotron 3.5 Lightning 30B-A3BNew

Mamba-2 + MoE hybrid built for the execution layer of agents: only 6 attention blocks, so the KV cache is almost free. Weights, data and recipe all open.

NVIDIA · 11 Aug 2026 · OpenMDW-1.1
Ling 3.0 Tiny 7.9B-A1.3BNew

An 8B MoE with 1.3B active and a latent KV cache on only 6 of 24 layers — reasoning and tool use sized for Apple Silicon and edge boxes.

inclusionAI · 10 Aug 2026 · MIT
Muse Glimmer 30BNew

Meta's first open weights since Llama 4: a dense 30B distilled from Muse Spark for always-on local agents. Two KV heads keep the cache small.

Meta · 10 Aug 2026 · Apache 2.0
Ling 3.0 Flash 124B-A5BNew

A 124B hybrid (5 linear-attention layers per MLA layer) with 5.1B active: SWE-bench Pro 56.6 and AIME 93 claimed. Built for 96–128 GB machines.

inclusionAI · 2 Aug 2026 · MIT
DeepSeek V4 Flash 284B-A13BNew

The V4 that 128 GB machines can actually run at Q3. Cache is modelled as a 576-wide latent; V4 compresses it further at long context, so this is conservative.

DeepSeek · 31 Jul 2026 · MIT
LFM2.5 2.6BNew

Convolution-heavy hybrid for CPUs and NPUs: 22 of 30 blocks keep no KV cache at all.

Liquid AI · 28 Jul 2026 · LFM Open License v1.0

91

GPUs & chips profiled

79

Models sized

16

New in the last month

5

Runtimes covered

computed

Not benchmarked

Start here

Pick the device, then the job

VRAM is the binding constraint, and it is the one number you already know. Choose your card and you get a recommendation per use case — best overall, coding, reasoning, vision, agents, fastest — plus every model in the catalogue sorted into runs-great, runs-tight and won't-fit.

Browse all 91 devices →

The trap

The weights are only half of it

A 32B model at Q4_K_M is 18.4 GB of weights, which looks fine on a 24 GB card — until you ask for 128K context and the KV cache alone wants 32 GB. Every table here prices the cache at the context you actually chose.

How the sizing works →

Then

One command, not a weekend

Every model page ends in the exact command for Ollama, llama.cpp, LM Studio or MLX, with the context flag already set to the size we said would fit.

Compare the four runtimes →

Not sure local is worth it yet? The honest comparison is cost per unit of work: a GPU is one payment, an API is a rate per million tokens. The AI article price calculator costs a finished article across 40 hosted Claude, GPT-5, Gemini and Grok models — hold that against the card you were about to buy.