Ollama models

Ollama models: which ones fit your GPU?

Updated 8 Oct 2026 1 new this month

62 of the 80 models in the catalogue are in the Ollama library. Each row is the tag to pull, the size of the weights Ollama downloads by default (its library tags are Q4_K_M unless the model ships a native format), and the smallest card that runs it with headroom at an 8K context. Pick a device to see the verdict and estimated speed on it.

Ollama's default context is 4K tokens, not the model's maximum. Set OLLAMA_CONTEXT_LENGTH (every model page here prints the line with the size that fits) or the model silently forgets the start of long prompts. The KV cache for that context is the part of the memory bill the download size does not show — the KV cache calculator prices it.

On this device any other device →
ModelPullParamsDownloadMax contextNeeds
Qwen3 0.6B
Useful mostly as a speculative-decoding draft model for its larger siblings.
ollama pull qwen3:0.6b 0.6B 0.3 GB 32K 8 GB
Qwen3.5 0.8B
Draft model for speculative decoding, or a classifier that fits in 1 GB.
ollama pull qwen3.5:0.8b 0.87B 0.5 GB 256K 8 GB
Gemma 3 1B
Text-only. Single KV head makes its cache almost free at long context.
ollama pull gemma3:1b 1.0B 0.6 GB 32K 8 GB
Llama 3.2 1B Instruct
The smallest Llama worth running. Fits anywhere, including phones and 4 GB cards.
ollama pull llama3.2:1b 1.24B 0.7 GB 128K 8 GB
Qwen3 1.7B
Punches above its size on structured tasks, with optional thinking mode.
ollama pull qwen3:1.7b 1.72B 1.0 GB 32K 8 GB
Qwen3.5 2B
Phone-class, and multimodal. Replaces Llama 3.2 3B as the "it runs on anything" answer.
ollama pull qwen3.5:2b 2.27B 1.3 GB 256K 8 GB
Llama 3.2 3B Instruct
The 2024 "it just runs" model for 8 GB laptops. Qwen3.5 4B does the same job better now.
ollama pull llama3.2:3b 3.21B 1.8 GB 128K 8 GB
Granite 4.1 3B
Dense, small, enterprise-flavoured: tool calling and instruction following, no thinking mode.
ollama pull granite4.1:3b 3.4B 1.9 GB 128K 8 GB
Phi-4-mini 3.8B
MIT-licensed, dense, and unusually strong on instruction following for its size.
ollama pull phi4-mini:3.8b 3.84B 2.2 GB 128K 8 GB
Ministral 3 3B
Edge model with a vision encoder and a 256K window. Apache 2.0.
ollama pull ministral-3:3b 3.85B 2.2 GB 256K 8 GB
Qwen3 4B
The 2025 sweet spot for 8 GB cards with reasoning traces. Qwen3.5 4B adds vision and 8× the context.
ollama pull qwen3:4b 4.02B 2.3 GB 32K 8 GB
Gemma 3 4B
Vision-capable at 4B. Superseded by Gemma 4 E4B, still everywhere.
ollama pull gemma3:4b 4.3B 2.4 GB 128K 8 GB
Qwen3.5 4B
The 8 GB coding agent. Q4 lands near 3.4 GB, leaving room for a real context window.
ollama pull qwen3.5:4b 4.66B 2.6 GB 256K 8 GB
Gemma 4 E2B
"E2B" is 2.3B effective, but the file holds 5B because of per-layer embeddings — size it as 5B. Text, image and audio in.
ollama pull gemma4:e2b 5.1B 2.9 GB 128K 8 GB
Mistral 7B Instruct v0.3
Old but extremely well behaved, and permissively licensed for commercial use.
ollama pull mistral:7b 7.25B 4.1 GB 32K 8 GB
Olmo 3 7B Instruct
Fully open — training data and code included. Full multi-head attention, so its cache is 4× a GQA 7B at the same context.
ollama pull olmo-3:7b 7.3B 4.1 GB 64K 10 GB
DeepSeek-R1-Distill-Qwen 7B
Reasoning traces on a 7B budget. Expect long outputs — budget context accordingly.
ollama pull deepseek-r1:7b 7.62B 4.3 GB 128K 8 GB
Qwen2.5-Coder 7B
The standard local autocomplete model — small enough to keep resident all day, and still the best FIM model under 8B.
ollama pull qwen2.5-coder:7b 7.62B 4.3 GB 128K 8 GB
Gemma 4 E4B
The laptop Gemma. 4.5B effective, 8B on disk; a single KV head per window layer keeps its cache tiny.
ollama pull gemma4:e4b 8.0B 4.5 GB 128K 8 GB
Llama 3.1 8B Instruct
Still the most widely deployed local model, with the largest fine-tune ecosystem. Not the strongest 8B any more.
ollama pull llama3.1:8b 8.03B 4.5 GB 128K 10 GB
Qwen3 8B
Apache-2.0 alternative to Llama 3.1 8B, with a switchable thinking mode.
ollama pull qwen3:8b 8.19B 4.6 GB 128K 10 GB
LFM2.5 8B-A1B
An 8B MoE with ~1.5B active, aimed at laptops without a GPU. Licence is permissive below $10M revenue.
ollama pull lfm2.5:8b 8.47B MoE 4.8 GB 125K 8 GB
Granite 4.1 8B
Matches the old Granite 4.0 32B MoE at a quarter of the size. Fast, Apache 2.0, no reasoning traces.
ollama pull granite4.1:8b 8.79B 4.9 GB 128K 10 GB
Ministral 3 8B
Mistral's 8B with images in. Plain GQA, so budget more KV cache than Qwen3.5 9B at the same context.
ollama pull ministral-3:8b 8.92B 5.0 GB 256K 10 GB
Ornith 1.5 9B
The small Ornith: a coding-agent reasoning build on the Qwen3.5 9B architecture. Same VRAM as its base, thinks before every answer.
ollama pull ornith-1.5:9b 9.41B 5.3 GB 256K 10 GB
Qwen3.5 9B
The default for 8–12 GB cards in 2026: beats every older 8B on every published benchmark, with vision.
ollama pull qwen3.5:9b 9.65B 5.4 GB 256K 10 GB
Gemma 4 12B
The "unified" Gemma 4: text, image and audio in one 12B that fits a 12 GB card at Q4. 140+ languages.
ollama pull gemma4:12b 12B 6.7 GB 256K 11 GB
Gemma 3 12B
Strong multilingual chat with images, sized for 12–16 GB cards. Gemma 4 12B is the same size and better.
ollama pull gemma3:12b 12.2B 6.9 GB 128K 12 GB
Mistral NeMo 12B
Multilingual 12B with a 128K window, built with NVIDIA. A roleplay and fiction favourite that refuses to die.
ollama pull mistral-nemo:12b 12.2B 6.9 GB 128K 12 GB
Ministral 3 14B
The largest Ministral. A 12 GB card runs it at Q4 with a few gigabytes to spare.
ollama pull ministral-3:14b 13.9B 7.8 GB 256K 16 GB
Phi-4 14B
Trained heavily on synthetic reasoning data. Short 16K window is its main limitation.
ollama pull phi4:14b 14.7B 8.3 GB 16K 16 GB
Qwen3 14B
The largest Qwen3 that fits a 12 GB card at Q4 with room for context.
ollama pull qwen3:14b 14.8B 8.3 GB 128K 16 GB
DeepSeek-R1-Distill-Qwen 14B
The 2025 reasoning-per-gigabyte pick for a 12 GB card. Qwen3.5 9B in thinking mode has since overtaken it.
ollama pull deepseek-r1:14b 14.8B 8.3 GB 128K 16 GB
Qwen2.5-Coder 14B
Noticeably better at whole-file edits than the 7B, still comfortable on 12 GB.
ollama pull qwen2.5-coder:14b 14.8B 8.3 GB 128K 16 GB
gpt-oss 20B
Ships natively in MXFP4, so the 4-bit weights are the reference weights, not a lossy copy. Fits 16 GB.
ollama pull gpt-oss:20b 20.9B MoE 10.8 GB 128K 16 GB
Mistral Small 3.2 24B
Apache-2.0, vision-capable, and the most 24 GB-friendly of the 2025 generalists.
ollama pull mistral-small:24b 23.6B 13.3 GB 128K 20 GB
Gemma 4 26B-A4B
Mixture of experts with 3.8B active. Slower to think than Qwen3.6 35B-A3B, faster to answer, and it sees images.
ollama pull gemma4:26b 26.5B MoE 14.9 GB 256K 24 GB
Gemma 3 27B
The 2025 single-GPU generalist with vision. Its Gemma-licence terms are the reason to prefer Gemma 4 now.
ollama pull gemma3:27b 27.4B 15.4 GB 128K 24 GB
Qwen3.8 27B
The current default local Qwen: dense 27B, text + image + video, 262K context. Only 16 of its 64 blocks keep a KV cache, so long context is cheap.
ollama pull qwen3.8:27b 27.8B 15.6 GB 256K 24 GB
Qwen3.6 27B
The 24 GB coding pick of spring 2026 (77.2 SWE-bench Verified). Same shape as 3.8, one generation behind.
ollama pull qwen3.6:27b 27.8B 15.6 GB 256K 24 GB
Granite 4.1 30B
The largest Granite. Dense 29B at Q4 is a comfortable 24 GB fit.
ollama pull granite4.1:30b 28.9B 16.3 GB 128K 24 GB
Muse Glimmer 30B
Meta's first open weights since Llama 4: a dense 30B distilled from Muse Spark for always-on local agents. Two KV heads keep the cache small.
ollama pull muse-glimmer:30b 29.8B 16.8 GB 128K 24 GB
Qwen3 30B-A3B
The MoE that made "3B active" a category. Qwen3.6 35B-A3B is its direct replacement.
ollama pull qwen3:30b-a3b 30.5B MoE 17.1 GB 128K 24 GB
Qwen3 Coder 30B-A3B
Agentic coding MoE with a 256K native window. Still the most-downloaded local code model.
ollama pull qwen3-coder:30b 30.5B MoE 17.1 GB 256K 24 GB
GLM-4.7-Flash 30B-A3B
MIT-licensed 30B-A3B tuned for agentic coding, with a DeepSeek-style latent KV cache. 60–80 tok/s reported on a 4090.
ollama pull glm-4.7-flash:latest 31.2B MoE 17.5 GB 198K 24 GB
Gemma 4 31B
The dense flagship: strongest maths of the 24–32 GB class (89% AIME), clean prose, vision. Q4 is a tight 24 GB fit.
ollama pull gemma4:31b 31.3B 17.6 GB 256K 32 GB
Nemotron 3.5 Lightning 30B-A3B
Mamba-2 + MoE hybrid built for the execution layer of agents: only 6 attention blocks, so the KV cache is almost free. Weights, data and recipe all open.
ollama pull nemotron-3.5-lightning:30b-a3b 31.6B MoE 17.8 GB 256K 24 GB
Olmo 3.1 32B Instruct
The largest fully open model you can audit end to end. Q4 fits 24 GB, tightly.
ollama pull olmo-3:32b 32.2B 18.1 GB 64K 32 GB
Qwen3 32B
The classic 24 GB target, and still the strongest local translator under 70B. Qwen3.8 27B is smaller and better at everything else.
ollama pull qwen3:32b 32.8B 18.4 GB 128K 32 GB
DeepSeek-R1-Distill-Qwen 32B
MIT-licensed and close to the 70B distill on maths. A 24 GB card handles it at Q4.
ollama pull deepseek-r1:32b 32.8B 18.4 GB 128K 32 GB
Qwen2.5-Coder 32B
The first local code model that felt competitive with hosted assistants. Qwen3.8 27B is smaller and far ahead on agentic work.
ollama pull qwen2.5-coder:32b 32.8B 18.4 GB 128K 32 GB
Qwen3.6 35B-A3B
Mixture of experts with ~3B active: the fastest serious model a 24 GB card runs, and the best MoE under 40B on agentic coding.
ollama pull qwen3.6:35b-a3b 35.9B MoE 20.2 GB 256K 32 GB
Ornith 1.5 35B-A3B
A reasoning-first MIT build on the Qwen3.6 35B-A3B architecture (thinks before every answer). Same VRAM as its base.
ollama pull ornith-1.5:35b 35.9B MoE 20.2 GB 256K 32 GB
Llama 3.3 70B Instruct
Still the creative-writing favourite: consistent voice, takes direction. Needs 48 GB to sit comfortably on GPU at Q4.
ollama pull llama3.3:70b 70.6B 39.7 GB 128K 64 GB
DeepSeek-R1-Distill-Llama 70B
The strongest of the R1 distills, and the one that most needs 48 GB or more.
ollama pull deepseek-r1:70b 70.6B 39.7 GB 128K 64 GB
Llama 4 Scout 109B-A17B
A 10M-token window on paper, chunked attention in practice (8K chunks on 3 of 4 layers). Needs 64 GB+ at Q4.
ollama pull llama4:scout 109B MoE 61.3 GB 1024K 80 GB
gpt-oss 120B
Designed to land on one 80 GB card. Only ~5B parameters are active per token.
ollama pull gpt-oss:120b 117B MoE 60.5 GB 128K 80 GB
Qwen3.5 122B-A10B
The 96–128 GB unified-memory model: 122B of knowledge at 10B-active speed.
ollama pull qwen3.5:122b-a10b 125B MoE 70.3 GB 256K 141 GB
Qwen3.8-Flash-Next 180B-A6B
The open preview of the Qwen4 architecture: a 125B-A6B hybrid (Gated DeltaNet + sparse attention, KV cache on 12 of 48 blocks) plus a 51B n-gram embedding and a 4B draft head — 180B on disk, 6B active. The hosted "Qwen3.8-Flash" is this model with a 1M window.
ollama pull qwen3.8-flash-next:125b-a6b-nvfp4 180B MoE 101.2 GB 256K 141 GB
Qwen3 235B-A22B
Workstation class. Realistically a 192 GB unified-memory or multi-GPU model.
ollama pull qwen3:235b-a22b 235B MoE 132.1 GB 128K 192 GB
Ornith 1.5 397B-A17B
The flagship Ornith on the Qwen3.5 397B-A17B architecture, MIT-licensed. A 256 GB Mac Studio at Q4, and it is in the Ollama library.
ollama pull ornith-1.5:397b 397B MoE 223.2 GB 256K multi-GPU
DeepSeek-R1 671B
The January 2025 moment. Multi-head latent attention keeps its KV cache tiny; the weights do not.
ollama pull deepseek-r1:671b 671B MoE 377.3 GB 128K multi-GPU

Models that are not in the Ollama library

18 catalogue models have no library tag — mostly the newest or largest releases. Ollama can still run most of them straight from Hugging Face with ollama run hf.co/ORG/MODEL-GGUF:Q4_K_M once a GGUF conversion exists; the models index links each one's model page, which sizes it and shows the llama.cpp command that always works.


Which quant is the tag?

A bare tag like qwen3.5:9b is Q4_K_M. Suffixes such as :9b-q8_0 pick another quantisation; what they mean.

Not sure what fits?

Best local LLM by VRAM class picks one model per job for 8 to 128 GB.

Ollama vs the rest

Five runtimes compared, and what LM Studio needs for the same models.