Ollama models

Ollama models: which ones fit your GPU?

Updated 8 Oct 2026 1 new this month

62 of the 80 models in the catalogue are in the Ollama library. Each row is the tag to pull, the size of the weights Ollama downloads by default (its library tags are Q4_K_M unless the model ships a native format), and the smallest card that runs it with headroom at an 8K context. Pick a device to see the verdict and estimated speed on it.

Ollama's default context is 4K tokens, not the model's maximum. Set OLLAMA_CONTEXT_LENGTH (every model page here prints the line with the size that fits) or the model silently forgets the start of long prompts. The KV cache for that context is the part of the memory bill the download size does not show — the KV cache calculator prices it.

On this device any other device →
ModelPullParamsDownloadMax contextNeedsOn Radeon RX 7900 XTX
Qwen3 0.6B
Useful mostly as a speculative-decoding draft model for its larger siblings.
ollama pull qwen3:0.6b 0.6B 0.3 GB 32K 8 GB 1723 tok/s
Qwen3.5 0.8B
Draft model for speculative decoding, or a classifier that fits in 1 GB.
ollama pull qwen3.5:0.8b 0.87B 0.5 GB 256K 8 GB 1188 tok/s
Gemma 3 1B
Text-only. Single KV head makes its cache almost free at long context.
ollama pull gemma3:1b 1.0B 0.6 GB 32K 8 GB 1034 tok/s
Llama 3.2 1B Instruct
The smallest Llama worth running. Fits anywhere, including phones and 4 GB cards.
ollama pull llama3.2:1b 1.24B 0.7 GB 128K 8 GB 834 tok/s
Qwen3 1.7B
Punches above its size on structured tasks, with optional thinking mode.
ollama pull qwen3:1.7b 1.72B 1.0 GB 32K 8 GB 601 tok/s
Qwen3.5 2B
Phone-class, and multimodal. Replaces Llama 3.2 3B as the "it runs on anything" answer.
ollama pull qwen3.5:2b 2.27B 1.3 GB 256K 8 GB 455 tok/s
Llama 3.2 3B Instruct
The 2024 "it just runs" model for 8 GB laptops. Qwen3.5 4B does the same job better now.
ollama pull llama3.2:3b 3.21B 1.8 GB 128K 8 GB 322 tok/s
Granite 4.1 3B
Dense, small, enterprise-flavoured: tool calling and instruction following, no thinking mode.
ollama pull granite4.1:3b 3.4B 1.9 GB 128K 8 GB 304 tok/s
Phi-4-mini 3.8B
MIT-licensed, dense, and unusually strong on instruction following for its size.
ollama pull phi4-mini:3.8b 3.84B 2.2 GB 128K 8 GB 269 tok/s
Ministral 3 3B
Edge model with a vision encoder and a 256K window. Apache 2.0.
ollama pull ministral-3:3b 3.85B 2.2 GB 256K 8 GB 268 tok/s
Qwen3 4B
The 2025 sweet spot for 8 GB cards with reasoning traces. Qwen3.5 4B adds vision and 8× the context.
ollama pull qwen3:4b 4.02B 2.3 GB 32K 8 GB 257 tok/s
Gemma 3 4B
Vision-capable at 4B. Superseded by Gemma 4 E4B, still everywhere.
ollama pull gemma3:4b 4.3B 2.4 GB 128K 8 GB 240 tok/s
Qwen3.5 4B
The 8 GB coding agent. Q4 lands near 3.4 GB, leaving room for a real context window.
ollama pull qwen3.5:4b 4.66B 2.6 GB 256K 8 GB 222 tok/s
Gemma 4 E2B
"E2B" is 2.3B effective, but the file holds 5B because of per-layer embeddings — size it as 5B. Text, image and audio in.
ollama pull gemma4:e2b 5.1B 2.9 GB 128K 8 GB 203 tok/s
Mistral 7B Instruct v0.3
Old but extremely well behaved, and permissively licensed for commercial use.
ollama pull mistral:7b 7.25B 4.1 GB 32K 8 GB 143 tok/s
Olmo 3 7B Instruct
Fully open — training data and code included. Full multi-head attention, so its cache is 4× a GQA 7B at the same context.
ollama pull olmo-3:7b 7.3B 4.1 GB 64K 10 GB 142 tok/s
DeepSeek-R1-Distill-Qwen 7B
Reasoning traces on a 7B budget. Expect long outputs — budget context accordingly.
ollama pull deepseek-r1:7b 7.62B 4.3 GB 128K 8 GB 136 tok/s
Qwen2.5-Coder 7B
The standard local autocomplete model — small enough to keep resident all day, and still the best FIM model under 8B.
ollama pull qwen2.5-coder:7b 7.62B 4.3 GB 128K 8 GB 136 tok/s
Gemma 4 E4B
The laptop Gemma. 4.5B effective, 8B on disk; a single KV head per window layer keeps its cache tiny.
ollama pull gemma4:e4b 8.0B 4.5 GB 128K 8 GB 129 tok/s
Llama 3.1 8B Instruct
Still the most widely deployed local model, with the largest fine-tune ecosystem. Not the strongest 8B any more.
ollama pull llama3.1:8b 8.03B 4.5 GB 128K 10 GB 129 tok/s
Qwen3 8B
Apache-2.0 alternative to Llama 3.1 8B, with a switchable thinking mode.
ollama pull qwen3:8b 8.19B 4.6 GB 128K 10 GB 126 tok/s
LFM2.5 8B-A1B
An 8B MoE with ~1.5B active, aimed at laptops without a GPU. Licence is permissive below $10M revenue.
ollama pull lfm2.5:8b 8.47B MoE 4.8 GB 125K 8 GB 265 tok/s
Granite 4.1 8B
Matches the old Granite 4.0 32B MoE at a quarter of the size. Fast, Apache 2.0, no reasoning traces.
ollama pull granite4.1:8b 8.79B 4.9 GB 128K 10 GB 118 tok/s
Ministral 3 8B
Mistral's 8B with images in. Plain GQA, so budget more KV cache than Qwen3.5 9B at the same context.
ollama pull ministral-3:8b 8.92B 5.0 GB 256K 10 GB 116 tok/s
Ornith 1.5 9B
The small Ornith: a coding-agent reasoning build on the Qwen3.5 9B architecture. Same VRAM as its base, thinks before every answer.
ollama pull ornith-1.5:9b 9.41B 5.3 GB 256K 10 GB 110 tok/s
Qwen3.5 9B
The default for 8–12 GB cards in 2026: beats every older 8B on every published benchmark, with vision.
ollama pull qwen3.5:9b 9.65B 5.4 GB 256K 10 GB 107 tok/s
Gemma 4 12B
The "unified" Gemma 4: text, image and audio in one 12B that fits a 12 GB card at Q4. 140+ languages.
ollama pull gemma4:12b 12B 6.7 GB 256K 11 GB 86 tok/s
Gemma 3 12B
Strong multilingual chat with images, sized for 12–16 GB cards. Gemma 4 12B is the same size and better.
ollama pull gemma3:12b 12.2B 6.9 GB 128K 12 GB 85 tok/s
Mistral NeMo 12B
Multilingual 12B with a 128K window, built with NVIDIA. A roleplay and fiction favourite that refuses to die.
ollama pull mistral-nemo:12b 12.2B 6.9 GB 128K 12 GB 85 tok/s
Ministral 3 14B
The largest Ministral. A 12 GB card runs it at Q4 with a few gigabytes to spare.
ollama pull ministral-3:14b 13.9B 7.8 GB 256K 16 GB 74 tok/s
Phi-4 14B
Trained heavily on synthetic reasoning data. Short 16K window is its main limitation.
ollama pull phi4:14b 14.7B 8.3 GB 16K 16 GB 70 tok/s
Qwen3 14B
The largest Qwen3 that fits a 12 GB card at Q4 with room for context.
ollama pull qwen3:14b 14.8B 8.3 GB 128K 16 GB 70 tok/s
DeepSeek-R1-Distill-Qwen 14B
The 2025 reasoning-per-gigabyte pick for a 12 GB card. Qwen3.5 9B in thinking mode has since overtaken it.
ollama pull deepseek-r1:14b 14.8B 8.3 GB 128K 16 GB 70 tok/s
Qwen2.5-Coder 14B
Noticeably better at whole-file edits than the 7B, still comfortable on 12 GB.
ollama pull qwen2.5-coder:14b 14.8B 8.3 GB 128K 16 GB 70 tok/s
gpt-oss 20B
Ships natively in MXFP4, so the 4-bit weights are the reference weights, not a lossy copy. Fits 16 GB.
ollama pull gpt-oss:20b 20.9B MoE 10.8 GB 128K 16 GB 120 tok/s
Mistral Small 3.2 24B
Apache-2.0, vision-capable, and the most 24 GB-friendly of the 2025 generalists.
ollama pull mistral-small:24b 23.6B 13.3 GB 128K 20 GB 44 tok/s
Gemma 4 26B-A4B
Mixture of experts with 3.8B active. Slower to think than Qwen3.6 35B-A3B, faster to answer, and it sees images.
ollama pull gemma4:26b 26.5B MoE 14.9 GB 256K 24 GB 105 tok/s
Gemma 3 27B
The 2025 single-GPU generalist with vision. Its Gemma-licence terms are the reason to prefer Gemma 4 now.
ollama pull gemma3:27b 27.4B 15.4 GB 128K 24 GB 38 tok/s
Qwen3.8 27B
The current default local Qwen: dense 27B, text + image + video, 262K context. Only 16 of its 64 blocks keep a KV cache, so long context is cheap.
ollama pull qwen3.8:27b 27.8B 15.6 GB 256K 24 GB 37 tok/s
Qwen3.6 27B
The 24 GB coding pick of spring 2026 (77.2 SWE-bench Verified). Same shape as 3.8, one generation behind.
ollama pull qwen3.6:27b 27.8B 15.6 GB 256K 24 GB 37 tok/s
Granite 4.1 30B
The largest Granite. Dense 29B at Q4 is a comfortable 24 GB fit.
ollama pull granite4.1:30b 28.9B 16.3 GB 128K 24 GB 36 tok/s
Muse Glimmer 30B
Meta's first open weights since Llama 4: a dense 30B distilled from Muse Spark for always-on local agents. Two KV heads keep the cache small.
ollama pull muse-glimmer:30b 29.8B 16.8 GB 128K 24 GB 35 tok/s
Qwen3 30B-A3B
The MoE that made "3B active" a category. Qwen3.6 35B-A3B is its direct replacement.
ollama pull qwen3:30b-a3b 30.5B MoE 17.1 GB 128K 24 GB 120 tok/s
Qwen3 Coder 30B-A3B
Agentic coding MoE with a 256K native window. Still the most-downloaded local code model.
ollama pull qwen3-coder:30b 30.5B MoE 17.1 GB 256K 24 GB 120 tok/s
GLM-4.7-Flash 30B-A3B
MIT-licensed 30B-A3B tuned for agentic coding, with a DeepSeek-style latent KV cache. 60–80 tok/s reported on a 4090.
ollama pull glm-4.7-flash:latest 31.2B MoE 17.5 GB 198K 24 GB 133 tok/s
Gemma 4 31B
The dense flagship: strongest maths of the 24–32 GB class (89% AIME), clean prose, vision. Q4 is a tight 24 GB fit.
ollama pull gemma4:31b 31.3B 17.6 GB 256K 32 GB 33 tok/s
Nemotron 3.5 Lightning 30B-A3B
Mamba-2 + MoE hybrid built for the execution layer of agents: only 6 attention blocks, so the KV cache is almost free. Weights, data and recipe all open.
ollama pull nemotron-3.5-lightning:30b-a3b 31.6B MoE 17.8 GB 256K 24 GB 124 tok/s
Olmo 3.1 32B Instruct
The largest fully open model you can audit end to end. Q4 fits 24 GB, tightly.
ollama pull olmo-3:32b 32.2B 18.1 GB 64K 32 GB 32 tok/s
Qwen3 32B
The classic 24 GB target, and still the strongest local translator under 70B. Qwen3.8 27B is smaller and better at everything else.
ollama pull qwen3:32b 32.8B 18.4 GB 128K 32 GB 32 tok/s
DeepSeek-R1-Distill-Qwen 32B
MIT-licensed and close to the 70B distill on maths. A 24 GB card handles it at Q4.
ollama pull deepseek-r1:32b 32.8B 18.4 GB 128K 32 GB 32 tok/s
Qwen2.5-Coder 32B
The first local code model that felt competitive with hosted assistants. Qwen3.8 27B is smaller and far ahead on agentic work.
ollama pull qwen2.5-coder:32b 32.8B 18.4 GB 128K 32 GB 32 tok/s
Qwen3.6 35B-A3B
Mixture of experts with ~3B active: the fastest serious model a 24 GB card runs, and the best MoE under 40B on agentic coding.
ollama pull qwen3.6:35b-a3b 35.9B MoE 20.2 GB 256K 32 GB 120 tok/s
Ornith 1.5 35B-A3B
A reasoning-first MIT build on the Qwen3.6 35B-A3B architecture (thinks before every answer). Same VRAM as its base.
ollama pull ornith-1.5:35b 35.9B MoE 20.2 GB 256K 32 GB 120 tok/s
Llama 3.3 70B Instruct
Still the creative-writing favourite: consistent voice, takes direction. Needs 48 GB to sit comfortably on GPU at Q4.
ollama pull llama3.3:70b 70.6B 39.7 GB 128K 64 GB offload
DeepSeek-R1-Distill-Llama 70B
The strongest of the R1 distills, and the one that most needs 48 GB or more.
ollama pull deepseek-r1:70b 70.6B 39.7 GB 128K 64 GB offload
Llama 4 Scout 109B-A17B
A 10M-token window on paper, chunked attention in practice (8K chunks on 3 of 4 layers). Needs 64 GB+ at Q4.
ollama pull llama4:scout 109B MoE 61.3 GB 1024K 80 GB won't fit
gpt-oss 120B
Designed to land on one 80 GB card. Only ~5B parameters are active per token.
ollama pull gpt-oss:120b 117B MoE 60.5 GB 128K 80 GB won't fit
Qwen3.5 122B-A10B
The 96–128 GB unified-memory model: 122B of knowledge at 10B-active speed.
ollama pull qwen3.5:122b-a10b 125B MoE 70.3 GB 256K 141 GB won't fit
Qwen3.8-Flash-Next 180B-A6B
The open preview of the Qwen4 architecture: a 125B-A6B hybrid (Gated DeltaNet + sparse attention, KV cache on 12 of 48 blocks) plus a 51B n-gram embedding and a 4B draft head — 180B on disk, 6B active. The hosted "Qwen3.8-Flash" is this model with a 1M window.
ollama pull qwen3.8-flash-next:125b-a6b-nvfp4 180B MoE 101.2 GB 256K 141 GB won't fit
Qwen3 235B-A22B
Workstation class. Realistically a 192 GB unified-memory or multi-GPU model.
ollama pull qwen3:235b-a22b 235B MoE 132.1 GB 128K 192 GB won't fit
Ornith 1.5 397B-A17B
The flagship Ornith on the Qwen3.5 397B-A17B architecture, MIT-licensed. A 256 GB Mac Studio at Q4, and it is in the Ollama library.
ollama pull ornith-1.5:397b 397B MoE 223.2 GB 256K multi-GPU won't fit
DeepSeek-R1 671B
The January 2025 moment. Multi-head latent attention keeps its KV cache tiny; the weights do not.
ollama pull deepseek-r1:671b 671B MoE 377.3 GB 128K multi-GPU won't fit

Models that are not in the Ollama library

18 catalogue models have no library tag — mostly the newest or largest releases. Ollama can still run most of them straight from Hugging Face with ollama run hf.co/ORG/MODEL-GGUF:Q4_K_M once a GGUF conversion exists; the models index links each one's model page, which sizes it and shows the llama.cpp command that always works.


Which quant is the tag?

A bare tag like qwen3.5:9b is Q4_K_M. Suffixes such as :9b-q8_0 pick another quantisation; what they mean.

Not sure what fits?

Best local LLM by VRAM class picks one model per job for 8 to 128 GB.

Ollama vs the rest

Five runtimes compared, and what LM Studio needs for the same models.