LM Studio · 2026
LM Studio hardware requirements: what you need to run each model
LM Studio itself runs on almost anything; the model you load is what has requirements. The app's own minimums are a few lines. What decides which models you can run — and how fast — is the GPU or unified memory underneath, so most of this page is that, computed from the same catalogue as the rest of the site.
The app's minimums
- macOS: Apple Silicon (M1 or later). Intel Macs are no longer supported.
- Windows: x64 or ARM64 (Snapdragon X). The CPU must support AVX2 — anything from the last decade does.
- Linux: x64, shipped as an AppImage, AVX2 required.
- RAM: 16 GB recommended. The app is light; the number is about holding a model when there is no GPU to hold it.
- GPU: optional. NVIDIA (CUDA), AMD (ROCm or Vulkan) and Intel Arc (Vulkan) are all driven; on a Mac the GPU is the unified memory.
These move with releases; lmstudio.ai has the current list. Everything below does not move, because it is arithmetic.
What LM Studio can run on your memory
LM Studio loads GGUF files through llama.cpp and, on a Mac, MLX files through Apple's framework, so its hardware requirements per model are the same as any other runtime's: the quantised weights, plus a KV cache for the context you set, plus about 0.6 GB of overhead, against what your GPU can actually use. By memory class, at the default Q4_K_M and an 8K context:
| Memory | Usable | Pick for most people | Largest that fits | Models fit | |
|---|---|---|---|---|---|
| 8 GB VRAM e.g. GeForce RTX 4060 |
7.0 GB | Gemma 4 E4B 5.2 GB | LFM2.5 8B-A1B Q4_K_M | 28 | All of them |
| 12 GB VRAM e.g. GeForce RTX 3060 12 GB |
10.6 GB | Gemma 4 12B 8.2 GB | Gemma 3 12B Q4_K_M | 37 | All of them |
| 16 GB VRAM e.g. GeForce RTX 5080 |
14.4 GB | Gemma 4 12B 8.2 GB | gpt-oss 20B MXFP4 | 38 | All of them |
| 24 GB VRAM e.g. GeForce RTX 4090 |
22.4 GB | Qwen3.8 27B 16.7 GB | Nemotron 3.5 Lightning 30B-A3B Q4_K_M | 58 | All of them |
| 32 GB VRAM e.g. GeForce RTX 5090 |
30.4 GB | Qwen3.8 27B 16.7 GB | Qwen3.6 35B-A3B Q4_K_M | 58 | All of them |
| 48 GB VRAM e.g. RTX 6000 Ada |
46.4 GB | Qwen3.8 27B 16.7 GB | Qwen3.6 35B-A3B Q4_K_M | 60 | All of them |
| 64 GB unified memory e.g. M4 Max · 64 GB |
48.0 GB | Qwen3.8 27B 16.7 GB | Qwen3.6 35B-A3B Q4_K_M | 60 | All of them |
| 128 GB unified memory e.g. M4 Max · 128 GB |
96.0 GB | Qwen3.8 27B 16.7 GB | Qwen3.5 122B-A10B Q4_K_M | 66 | All of them |
A 24 GB card is the practical sweet spot for the current 27B generation; a 64 GB Mac is the cheapest way to a 70B. Not sure which class you are in? Detect your machine, or pick it from the catalogue and choose LM Studio as the runtime on any model page.
Requirements for popular models
Download size at the quantisation LM Studio suggests by default, the smallest GPU that runs it with 15% headroom at 8K, and the smallest Mac that does.
| Model | Download | Smallest GPU | Smallest Mac | Max context |
|---|---|---|---|---|
| Qwen3.5 4B 4.66B |
2.6 GB Q4_K_M | 8 GB GPU | 16 GB Mac | 256K |
| Qwen3.5 9B 9.65B |
5.4 GB Q4_K_M | 10 GB GPU | 16 GB Mac | 256K |
| Gemma 4 12B 12B |
6.7 GB Q4_K_M | 11 GB GPU | 16 GB Mac | 256K |
| gpt-oss 20B 20.9B MoE |
10.8 GB MXFP4 | 16 GB GPU | 24 GB Mac | 128K |
| Mistral Small 3.2 24B 23.6B |
13.3 GB Q4_K_M | 20 GB GPU | 24 GB Mac | 128K |
| Qwen3.8 27B 27.8B |
15.6 GB Q4_K_M | 24 GB GPU | 32 GB Mac | 256K |
| Gemma 4 31B 31.3B |
17.6 GB Q4_K_M | 32 GB GPU | 32 GB Mac | 256K |
| Qwen3.6 35B-A3B 35.9B MoE |
20.2 GB Q4_K_M | 32 GB GPU | 36 GB Mac | 256K |
| Llama 3.3 70B Instruct 70.6B |
39.7 GB Q4_K_M | 64 GB GPU | 96 GB Mac | 128K |
| gpt-oss 120B 117B MoE |
60.5 GB MXFP4 | 80 GB GPU | 128 GB Mac | 128K |
| GLM-5.3-Flash 320B-A18B 321B MoE |
180.5 GB Q4_K_M | no single GPU | 512 GB Mac | 1024K |
| DeepSeek V4.1 Flash 552B-A16B 552B MoE |
310.4 GB Q4_K_M | no single GPU | 512 GB Mac | 1024K |
The settings that change the requirements
LM Studio shows a fit estimate — "full GPU offload possible", "partial GPU offload possible", or a warning — before you download, using the same quantities as this site. Four settings in the load dialog move the answer:
- Quantisation. The model search lists one file per quant. Q4_K_M is the default and the right one; step down to Q3_K_M only to afford a longer context, step up to Q6_K or Q8_0 when memory is spare. What the letters mean.
- Context length. It defaults to a modest window, not the model's maximum. Every token you add costs KV cache; for a dense 27B that is tens of gigabytes by 128K. Set it to the size the model page here says fits, no higher. Price it.
- GPU offload. The slider is the number of layers on the GPU. "Max" is what you want when the model fits; when it does not, every layer left on the CPU streams from system RAM at a few tokens per second, which is why a partial-offload "possible" is not a recommendation.
- KV cache quantisation and flash attention. A q8_0 cache roughly halves the context cost for a negligible loss, and flash attention has to be on for it. On a tight fit this is the first thing to try.
The command-line route
The lms CLI drives the same engine without the window. A 9B assistant on a 16 GB card, with the context set to the size that fits and every layer on the GPU:
$ lms get Qwen/Qwen3.5-9B $ lms load qwen3.5-9b --context-length 8192 \ --gpu max
The same model on a 48 GB Mac, where LM Studio will pick MLX if there is an MLX conversion and llama.cpp otherwise:
$ lms get Qwen/Qwen3.8-27B $ lms load qwen3.8-27b --context-length 8192 \ --gpu max
Every model page on this site prints the lms command for the device you chose, with the context and offload already set to what fits — pick LM Studio in the runtime selector.
Without a GPU
LM Studio runs on the CPU alone through llama.cpp. Speed is RAM bandwidth divided by model size, so a DDR5 desktop manages a 4B model at a usable pace and mixture-of-experts models with a few billion active parameters run surprisingly well. At 5 tokens per second or better on a CPU only · DDR5 dual-channel: Qwen3.6 35B-A3B (9.0), Ornith 1.5 35B-A3B (9.0), Nemotron 3.5 Lightning 30B-A3B (9.3), GLM-4.7-Flash 30B-A3B (9.9). The 16 GB RAM recommendation is about this case: it holds a 9B at Q4 with room for the OS.
LM Studio versus Ollama
Same engine, different defaults. LM Studio shows the fit before downloading, lets you pick the exact quantisation file and exposes the offload and cache settings in the UI; Ollama hides all four behind a tag and a Modelfile, and defaults the context to 4K. If you would rather click than type, LM Studio is the better default; if you want a background service with an API, Ollama is. The five runtimes compared · Ollama models by VRAM.