Ollama models
Ollama models: which ones fit your GPU?
62 of the 80 models in the catalogue are in the Ollama library. Each row is the tag to pull, the size of the weights Ollama downloads by default (its library tags are Q4_K_M unless the model ships a native format), and the smallest card that runs it with headroom at an 8K context. Pick a device to see the verdict and estimated speed on it.
Ollama's default context is 4K tokens, not the model's maximum. Set OLLAMA_CONTEXT_LENGTH (every model page here prints the
line with the size that fits) or the model silently forgets the start of long prompts. The KV cache for that context is the part of the memory bill the
download size does not show — the KV cache calculator prices it.
| Model | Pull | Params | Download | Max context | Needs |
|---|---|---|---|---|---|
| Qwen3 0.6B Useful mostly as a speculative-decoding draft model for its larger siblings. |
ollama pull qwen3:0.6b | 0.6B | 0.3 GB | 32K | 8 GB |
| Qwen3.5 0.8B Draft model for speculative decoding, or a classifier that fits in 1 GB. |
ollama pull qwen3.5:0.8b | 0.87B | 0.5 GB | 256K | 8 GB |
| Gemma 3 1B Text-only. Single KV head makes its cache almost free at long context. |
ollama pull gemma3:1b | 1.0B | 0.6 GB | 32K | 8 GB |
| Llama 3.2 1B Instruct The smallest Llama worth running. Fits anywhere, including phones and 4 GB cards. |
ollama pull llama3.2:1b | 1.24B | 0.7 GB | 128K | 8 GB |
| Qwen3 1.7B Punches above its size on structured tasks, with optional thinking mode. |
ollama pull qwen3:1.7b | 1.72B | 1.0 GB | 32K | 8 GB |
| Qwen3.5 2B Phone-class, and multimodal. Replaces Llama 3.2 3B as the "it runs on anything" answer. |
ollama pull qwen3.5:2b | 2.27B | 1.3 GB | 256K | 8 GB |
| Llama 3.2 3B Instruct The 2024 "it just runs" model for 8 GB laptops. Qwen3.5 4B does the same job better now. |
ollama pull llama3.2:3b | 3.21B | 1.8 GB | 128K | 8 GB |
| Granite 4.1 3B Dense, small, enterprise-flavoured: tool calling and instruction following, no thinking mode. |
ollama pull granite4.1:3b | 3.4B | 1.9 GB | 128K | 8 GB |
| Phi-4-mini 3.8B MIT-licensed, dense, and unusually strong on instruction following for its size. |
ollama pull phi4-mini:3.8b | 3.84B | 2.2 GB | 128K | 8 GB |
| Ministral 3 3B Edge model with a vision encoder and a 256K window. Apache 2.0. |
ollama pull ministral-3:3b | 3.85B | 2.2 GB | 256K | 8 GB |
| Qwen3 4B The 2025 sweet spot for 8 GB cards with reasoning traces. Qwen3.5 4B adds vision and 8× the context. |
ollama pull qwen3:4b | 4.02B | 2.3 GB | 32K | 8 GB |
| Gemma 3 4B Vision-capable at 4B. Superseded by Gemma 4 E4B, still everywhere. |
ollama pull gemma3:4b | 4.3B | 2.4 GB | 128K | 8 GB |
| Qwen3.5 4B The 8 GB coding agent. Q4 lands near 3.4 GB, leaving room for a real context window. |
ollama pull qwen3.5:4b | 4.66B | 2.6 GB | 256K | 8 GB |
| Gemma 4 E2B "E2B" is 2.3B effective, but the file holds 5B because of per-layer embeddings — size it as 5B. Text, image and audio in. |
ollama pull gemma4:e2b | 5.1B | 2.9 GB | 128K | 8 GB |
| Mistral 7B Instruct v0.3 Old but extremely well behaved, and permissively licensed for commercial use. |
ollama pull mistral:7b | 7.25B | 4.1 GB | 32K | 8 GB |
| Olmo 3 7B Instruct Fully open — training data and code included. Full multi-head attention, so its cache is 4× a GQA 7B at the same context. |
ollama pull olmo-3:7b | 7.3B | 4.1 GB | 64K | 10 GB |
| DeepSeek-R1-Distill-Qwen 7B Reasoning traces on a 7B budget. Expect long outputs — budget context accordingly. |
ollama pull deepseek-r1:7b | 7.62B | 4.3 GB | 128K | 8 GB |
| Qwen2.5-Coder 7B The standard local autocomplete model — small enough to keep resident all day, and still the best FIM model under 8B. |
ollama pull qwen2.5-coder:7b | 7.62B | 4.3 GB | 128K | 8 GB |
| Gemma 4 E4B The laptop Gemma. 4.5B effective, 8B on disk; a single KV head per window layer keeps its cache tiny. |
ollama pull gemma4:e4b | 8.0B | 4.5 GB | 128K | 8 GB |
| Llama 3.1 8B Instruct Still the most widely deployed local model, with the largest fine-tune ecosystem. Not the strongest 8B any more. |
ollama pull llama3.1:8b | 8.03B | 4.5 GB | 128K | 10 GB |
| Qwen3 8B Apache-2.0 alternative to Llama 3.1 8B, with a switchable thinking mode. |
ollama pull qwen3:8b | 8.19B | 4.6 GB | 128K | 10 GB |
| LFM2.5 8B-A1B An 8B MoE with ~1.5B active, aimed at laptops without a GPU. Licence is permissive below $10M revenue. |
ollama pull lfm2.5:8b | 8.47B MoE | 4.8 GB | 125K | 8 GB |
| Granite 4.1 8B Matches the old Granite 4.0 32B MoE at a quarter of the size. Fast, Apache 2.0, no reasoning traces. |
ollama pull granite4.1:8b | 8.79B | 4.9 GB | 128K | 10 GB |
| Ministral 3 8B Mistral's 8B with images in. Plain GQA, so budget more KV cache than Qwen3.5 9B at the same context. |
ollama pull ministral-3:8b | 8.92B | 5.0 GB | 256K | 10 GB |
| Ornith 1.5 9B The small Ornith: a coding-agent reasoning build on the Qwen3.5 9B architecture. Same VRAM as its base, thinks before every answer. |
ollama pull ornith-1.5:9b | 9.41B | 5.3 GB | 256K | 10 GB |
| Qwen3.5 9B The default for 8–12 GB cards in 2026: beats every older 8B on every published benchmark, with vision. |
ollama pull qwen3.5:9b | 9.65B | 5.4 GB | 256K | 10 GB |
| Gemma 4 12B The "unified" Gemma 4: text, image and audio in one 12B that fits a 12 GB card at Q4. 140+ languages. |
ollama pull gemma4:12b | 12B | 6.7 GB | 256K | 11 GB |
| Gemma 3 12B Strong multilingual chat with images, sized for 12–16 GB cards. Gemma 4 12B is the same size and better. |
ollama pull gemma3:12b | 12.2B | 6.9 GB | 128K | 12 GB |
| Mistral NeMo 12B Multilingual 12B with a 128K window, built with NVIDIA. A roleplay and fiction favourite that refuses to die. |
ollama pull mistral-nemo:12b | 12.2B | 6.9 GB | 128K | 12 GB |
| Ministral 3 14B The largest Ministral. A 12 GB card runs it at Q4 with a few gigabytes to spare. |
ollama pull ministral-3:14b | 13.9B | 7.8 GB | 256K | 16 GB |
| Phi-4 14B Trained heavily on synthetic reasoning data. Short 16K window is its main limitation. |
ollama pull phi4:14b | 14.7B | 8.3 GB | 16K | 16 GB |
| Qwen3 14B The largest Qwen3 that fits a 12 GB card at Q4 with room for context. |
ollama pull qwen3:14b | 14.8B | 8.3 GB | 128K | 16 GB |
| DeepSeek-R1-Distill-Qwen 14B The 2025 reasoning-per-gigabyte pick for a 12 GB card. Qwen3.5 9B in thinking mode has since overtaken it. |
ollama pull deepseek-r1:14b | 14.8B | 8.3 GB | 128K | 16 GB |
| Qwen2.5-Coder 14B Noticeably better at whole-file edits than the 7B, still comfortable on 12 GB. |
ollama pull qwen2.5-coder:14b | 14.8B | 8.3 GB | 128K | 16 GB |
| gpt-oss 20B Ships natively in MXFP4, so the 4-bit weights are the reference weights, not a lossy copy. Fits 16 GB. |
ollama pull gpt-oss:20b | 20.9B MoE | 10.8 GB | 128K | 16 GB |
| Mistral Small 3.2 24B Apache-2.0, vision-capable, and the most 24 GB-friendly of the 2025 generalists. |
ollama pull mistral-small:24b | 23.6B | 13.3 GB | 128K | 20 GB |
| Gemma 4 26B-A4B Mixture of experts with 3.8B active. Slower to think than Qwen3.6 35B-A3B, faster to answer, and it sees images. |
ollama pull gemma4:26b | 26.5B MoE | 14.9 GB | 256K | 24 GB |
| Gemma 3 27B The 2025 single-GPU generalist with vision. Its Gemma-licence terms are the reason to prefer Gemma 4 now. |
ollama pull gemma3:27b | 27.4B | 15.4 GB | 128K | 24 GB |
| Qwen3.8 27B The current default local Qwen: dense 27B, text + image + video, 262K context. Only 16 of its 64 blocks keep a KV cache, so long context is cheap. |
ollama pull qwen3.8:27b | 27.8B | 15.6 GB | 256K | 24 GB |
| Qwen3.6 27B The 24 GB coding pick of spring 2026 (77.2 SWE-bench Verified). Same shape as 3.8, one generation behind. |
ollama pull qwen3.6:27b | 27.8B | 15.6 GB | 256K | 24 GB |
| Granite 4.1 30B The largest Granite. Dense 29B at Q4 is a comfortable 24 GB fit. |
ollama pull granite4.1:30b | 28.9B | 16.3 GB | 128K | 24 GB |
| Muse Glimmer 30B Meta's first open weights since Llama 4: a dense 30B distilled from Muse Spark for always-on local agents. Two KV heads keep the cache small. |
ollama pull muse-glimmer:30b | 29.8B | 16.8 GB | 128K | 24 GB |
| Qwen3 30B-A3B The MoE that made "3B active" a category. Qwen3.6 35B-A3B is its direct replacement. |
ollama pull qwen3:30b-a3b | 30.5B MoE | 17.1 GB | 128K | 24 GB |
| Qwen3 Coder 30B-A3B Agentic coding MoE with a 256K native window. Still the most-downloaded local code model. |
ollama pull qwen3-coder:30b | 30.5B MoE | 17.1 GB | 256K | 24 GB |
| GLM-4.7-Flash 30B-A3B MIT-licensed 30B-A3B tuned for agentic coding, with a DeepSeek-style latent KV cache. 60–80 tok/s reported on a 4090. |
ollama pull glm-4.7-flash:latest | 31.2B MoE | 17.5 GB | 198K | 24 GB |
| Gemma 4 31B The dense flagship: strongest maths of the 24–32 GB class (89% AIME), clean prose, vision. Q4 is a tight 24 GB fit. |
ollama pull gemma4:31b | 31.3B | 17.6 GB | 256K | 32 GB |
| Nemotron 3.5 Lightning 30B-A3B Mamba-2 + MoE hybrid built for the execution layer of agents: only 6 attention blocks, so the KV cache is almost free. Weights, data and recipe all open. |
ollama pull nemotron-3.5-lightning:30b-a3b | 31.6B MoE | 17.8 GB | 256K | 24 GB |
| Olmo 3.1 32B Instruct The largest fully open model you can audit end to end. Q4 fits 24 GB, tightly. |
ollama pull olmo-3:32b | 32.2B | 18.1 GB | 64K | 32 GB |
| Qwen3 32B The classic 24 GB target, and still the strongest local translator under 70B. Qwen3.8 27B is smaller and better at everything else. |
ollama pull qwen3:32b | 32.8B | 18.4 GB | 128K | 32 GB |
| DeepSeek-R1-Distill-Qwen 32B MIT-licensed and close to the 70B distill on maths. A 24 GB card handles it at Q4. |
ollama pull deepseek-r1:32b | 32.8B | 18.4 GB | 128K | 32 GB |
| Qwen2.5-Coder 32B The first local code model that felt competitive with hosted assistants. Qwen3.8 27B is smaller and far ahead on agentic work. |
ollama pull qwen2.5-coder:32b | 32.8B | 18.4 GB | 128K | 32 GB |
| Qwen3.6 35B-A3B Mixture of experts with ~3B active: the fastest serious model a 24 GB card runs, and the best MoE under 40B on agentic coding. |
ollama pull qwen3.6:35b-a3b | 35.9B MoE | 20.2 GB | 256K | 32 GB |
| Ornith 1.5 35B-A3B A reasoning-first MIT build on the Qwen3.6 35B-A3B architecture (thinks before every answer). Same VRAM as its base. |
ollama pull ornith-1.5:35b | 35.9B MoE | 20.2 GB | 256K | 32 GB |
| Llama 3.3 70B Instruct Still the creative-writing favourite: consistent voice, takes direction. Needs 48 GB to sit comfortably on GPU at Q4. |
ollama pull llama3.3:70b | 70.6B | 39.7 GB | 128K | 64 GB |
| DeepSeek-R1-Distill-Llama 70B The strongest of the R1 distills, and the one that most needs 48 GB or more. |
ollama pull deepseek-r1:70b | 70.6B | 39.7 GB | 128K | 64 GB |
| Llama 4 Scout 109B-A17B A 10M-token window on paper, chunked attention in practice (8K chunks on 3 of 4 layers). Needs 64 GB+ at Q4. |
ollama pull llama4:scout | 109B MoE | 61.3 GB | 1024K | 80 GB |
| gpt-oss 120B Designed to land on one 80 GB card. Only ~5B parameters are active per token. |
ollama pull gpt-oss:120b | 117B MoE | 60.5 GB | 128K | 80 GB |
| Qwen3.5 122B-A10B The 96–128 GB unified-memory model: 122B of knowledge at 10B-active speed. |
ollama pull qwen3.5:122b-a10b | 125B MoE | 70.3 GB | 256K | 141 GB |
| Qwen3.8-Flash-Next 180B-A6B The open preview of the Qwen4 architecture: a 125B-A6B hybrid (Gated DeltaNet + sparse attention, KV cache on 12 of 48 blocks) plus a 51B n-gram embedding and a 4B draft head — 180B on disk, 6B active. The hosted "Qwen3.8-Flash" is this model with a 1M window. |
ollama pull qwen3.8-flash-next:125b-a6b-nvfp4 | 180B MoE | 101.2 GB | 256K | 141 GB |
| Qwen3 235B-A22B Workstation class. Realistically a 192 GB unified-memory or multi-GPU model. |
ollama pull qwen3:235b-a22b | 235B MoE | 132.1 GB | 128K | 192 GB |
| Ornith 1.5 397B-A17B The flagship Ornith on the Qwen3.5 397B-A17B architecture, MIT-licensed. A 256 GB Mac Studio at Q4, and it is in the Ollama library. |
ollama pull ornith-1.5:397b | 397B MoE | 223.2 GB | 256K | multi-GPU |
| DeepSeek-R1 671B The January 2025 moment. Multi-head latent attention keeps its KV cache tiny; the weights do not. |
ollama pull deepseek-r1:671b | 671B MoE | 377.3 GB | 128K | multi-GPU |
Models that are not in the Ollama library
18 catalogue models have no library tag — mostly the newest or largest releases.
Ollama can still run most of them straight from Hugging Face with ollama run hf.co/ORG/MODEL-GGUF:Q4_K_M once a GGUF conversion exists;
the models index links each one's model page, which sizes it and shows the llama.cpp command that always works.
Which quant is the tag?
A bare tag like qwen3.5:9b is Q4_K_M. Suffixes such as :9b-q8_0 pick another quantisation; what they mean.
Not sure what fits?
Best local LLM by VRAM class picks one model per job for 8 to 128 GB.
Ollama vs the rest
Five runtimes compared, and what LM Studio needs for the same models.