Best GPU for local LLMs · 2026
The best GPU for local LLMs, ranked by what it can actually run
91 GPUs, Apple Silicon chips and CPU configurations, ranked inside each memory class by what fits and how fast it runs. No prices — they move weekly and vary by country — and no benchmarks: the speeds are estimated from each card's published memory bandwidth, the same way for every device, so they are fair to compare and wrong in absolute terms by a similar factor.
1 · VRAM decides what runs
A model either fits in memory or it does not, and when it does not you lose an order of magnitude in speed to offload. Buy capacity first: 24 GB unlocks the current 27B generation, 16 GB the 20B MoEs, 12 GB a 12B with headroom, 8 GB a 9B tightly.
2 · Bandwidth decides how fast
Generation reads every active weight once per token, so tokens per second is bandwidth divided by model size. Within a class, that is the whole ranking — an RTX 3090's 936 GB/s runs a 27B at nearly the speed of a 4090's 1,008.
3 · Then the software
NVIDIA runs everything on day one. AMD and Intel run llama.cpp and Ollama well via Vulkan and ROCm, and vLLM on AMD only. Apple Silicon has MLX, which is faster than GGUF on the same Mac but lags on the newest releases.
Best 8 GB GPU for local LLMs
what to run on 8 GB →Enough for a 4–9B assistant and the small MoEs. A 9B at Q4 is a tight fit; 7B and under run with headroom.
| Device | Memory | Bandwidth | Models fit | Qwen3.5 9B tok/s | Qwen3.8 27B tok/s | Llama 3.3 70B Instruct tok/s | Best overall pick |
|---|---|---|---|---|---|---|---|
| GeForce RTX 3070 Ti fastest in class Ampere · GDDR6X · 2021 |
8 GB | 608 GB/s | 28 | 68 | ~3.5 offload | — | Gemma 4 E4B |
| Arc A750 Alchemist · GDDR6 · 2022 |
8 GB | 512 GB/s | 28 | 57 | ~3.5 offload | — | Gemma 4 E4B |
| GeForce RTX 5060 Blackwell · GDDR7 · 2025 |
8 GB | 448 GB/s | 28 | 50 | ~3.5 offload | — | Gemma 4 E4B |
| GeForce RTX 3070 Ampere · GDDR6 · 2020 |
8 GB | 448 GB/s | 28 | 50 | ~3.5 offload | — | Gemma 4 E4B |
| GeForce RTX 3060 Ti Ampere · GDDR6 · 2020 |
8 GB | 448 GB/s | 28 | 50 | ~3.5 offload | — | Gemma 4 E4B |
| GeForce RTX 4060 Ti Ada Lovelace · GDDR6 · 2023 |
8 GB | 288 GB/s | 28 | 32 | ~3.3 offload | — | Gemma 4 E4B |
| GeForce RTX 4060 Ada Lovelace · GDDR6 · 2023 |
8 GB | 272 GB/s | 28 | 30 | ~3.3 offload | — | Gemma 4 E4B |
| GeForce RTX 4070 Laptop Ada Lovelace · GDDR6 · 2023 |
8 GB | 256 GB/s | 28 | 29 | ~3.3 offload | — | Gemma 4 E4B |
Best 10–12 GB GPU for local LLMs
what to run on 10-12 GB →The first class where a 12B runs with headroom. Used RTX 3060 12 GB cards made this the default budget class.
| Device | Memory | Bandwidth | Models fit | Qwen3.5 9B tok/s | Qwen3.8 27B tok/s | Llama 3.3 70B Instruct tok/s | Best overall pick |
|---|---|---|---|---|---|---|---|
| GeForce RTX 3080 Ti fastest in class Ampere · GDDR6X · 2021 |
12 GB | 912 GB/s | 37 | 102 | ~5.4 offload | — | Gemma 4 12B |
| GeForce RTX 3080 12 GB Ampere · GDDR6X · 2022 |
12 GB | 912 GB/s | 37 | 102 | ~5.4 offload | — | Gemma 4 12B |
| GeForce RTX 5070 Blackwell · GDDR7 · 2025 |
12 GB | 672 GB/s | 37 | 75 | ~5.2 offload | — | Gemma 4 12B |
| GeForce RTX 4070 Ti Ada Lovelace · GDDR6X · 2023 |
12 GB | 504 GB/s | 37 | 56 | ~5.0 offload | — | Gemma 4 12B |
| GeForce RTX 4070 Super Ada Lovelace · GDDR6X · 2024 |
12 GB | 504 GB/s | 37 | 56 | ~5.0 offload | — | Gemma 4 12B |
| GeForce RTX 4070 Ada Lovelace · GDDR6X · 2023 |
12 GB | 504 GB/s | 37 | 56 | ~5.0 offload | — | Gemma 4 12B |
| Arc B580 Battlemage · GDDR6 · 2024 |
12 GB | 456 GB/s | 37 | 51 | ~4.9 offload | — | Gemma 4 12B |
| GeForce RTX 4080 Laptop Ada Lovelace · GDDR6 · 2023 |
12 GB | 432 GB/s | 37 | 48 | ~4.9 offload | — | Gemma 4 12B |
| Radeon RX 7700 XT RDNA 3 · GDDR6 · 2023 |
12 GB | 432 GB/s | 37 | 48 | ~4.9 offload | — | Gemma 4 12B |
| Radeon RX 6700 XT RDNA 2 · GDDR6 · 2021 |
12 GB | 384 GB/s | 37 | 43 | ~4.8 offload | — | Gemma 4 12B |
| GeForce RTX 3060 12 GB Ampere · GDDR6 · 2021 |
12 GB | 360 GB/s | 37 | 40 | ~4.7 offload | — | Gemma 4 12B |
| GeForce RTX 2060 12 GB Turing · GDDR6 · 2021 |
12 GB | 336 GB/s | 37 | 37 | ~4.6 offload | — | Gemma 4 12B |
| GeForce RTX 2080 Ti Turing · GDDR6 · 2018 |
11 GB | 616 GB/s | 33 | 69 | ~4.6 offload | — | Gemma 4 12B |
| GeForce GTX 1080 Ti Pascal · GDDR5X · 2017 |
11 GB | 484 GB/s | 33 | 54 | ~4.5 offload | — | Gemma 4 12B |
| GeForce RTX 3080 Ampere · GDDR6X · 2020 |
10 GB | 760 GB/s | 32 | 85 | ~4.3 offload | — | Qwen3.5 9B |
| Arc B570 Battlemage · GDDR6 · 2025 |
10 GB | 380 GB/s | 32 | 42 | ~4.0 offload | — | Qwen3.5 9B |
Best 16 GB GPU for local LLMs
what to run on 16 GB →gpt-oss 20B with its full 128K window, 14B dense models with room to spare, 24B tightly.
| Device | Memory | Bandwidth | Models fit | Qwen3.5 9B tok/s | Qwen3.8 27B tok/s | Llama 3.3 70B Instruct tok/s | Best overall pick |
|---|---|---|---|---|---|---|---|
| GeForce RTX 5080 fastest in class Blackwell · GDDR7 · 2025 |
16 GB | 960 GB/s | 38 | 107 | ~11 offload | — | Gemma 4 12B |
| GeForce RTX 5070 Ti Blackwell · GDDR7 · 2025 |
16 GB | 896 GB/s | 38 | 100 | ~11 offload | — | Gemma 4 12B |
| GeForce RTX 4080 Super Ada Lovelace · GDDR6X · 2024 |
16 GB | 736 GB/s | 38 | 82 | ~11 offload | — | Gemma 4 12B |
| GeForce RTX 4080 Ada Lovelace · GDDR6X · 2022 |
16 GB | 717 GB/s | 38 | 80 | ~11 offload | — | Gemma 4 12B |
| GeForce RTX 4070 Ti Super Ada Lovelace · GDDR6X · 2024 |
16 GB | 672 GB/s | 38 | 75 | ~10 offload | — | Gemma 4 12B |
| Radeon RX 9070 XT RDNA 4 · GDDR6 · 2025 |
16 GB | 645 GB/s | 38 | 72 | ~10 offload | — | Gemma 4 12B |
| Radeon RX 9070 RDNA 4 · GDDR6 · 2025 |
16 GB | 645 GB/s | 38 | 72 | ~10 offload | — | Gemma 4 12B |
| Radeon RX 7800 XT RDNA 3 · GDDR6 · 2023 |
16 GB | 624 GB/s | 38 | 70 | ~10 offload | — | Gemma 4 12B |
| GeForce RTX 4090 Laptop Ada Lovelace · GDDR6 · 2023 |
16 GB | 576 GB/s | 38 | 64 | ~9.8 offload | — | Gemma 4 12B |
| Radeon RX 7900 GRE RDNA 3 · GDDR6 · 2023 |
16 GB | 576 GB/s | 38 | 64 | ~9.8 offload | — | Gemma 4 12B |
| Arc A770 16 GB Alchemist · GDDR6 · 2022 |
16 GB | 560 GB/s | 38 | 62 | ~9.7 offload | — | Gemma 4 12B |
| Radeon RX 6900 XT RDNA 2 · GDDR6 · 2020 |
16 GB | 512 GB/s | 38 | 57 | ~9.3 offload | — | Gemma 4 12B |
| Radeon RX 6800 XT RDNA 2 · GDDR6 · 2020 |
16 GB | 512 GB/s | 38 | 57 | ~9.3 offload | — | Gemma 4 12B |
| GeForce RTX 5060 Ti 16 GB Blackwell · GDDR7 · 2025 |
16 GB | 448 GB/s | 38 | 50 | ~8.8 offload | — | Gemma 4 12B |
| RTX A4000 Ampere · GDDR6 ECC · 2021 |
16 GB | 448 GB/s | 38 | 50 | ~8.8 offload | — | Gemma 4 12B |
| GeForce RTX 4060 Ti 16 GB Ada Lovelace · GDDR6 · 2023 |
16 GB | 288 GB/s | 38 | 32 | ~7.1 offload | — | Gemma 4 12B |
| Radeon RX 7600 XT RDNA 3 · GDDR6 · 2024 |
16 GB | 288 GB/s | 38 | 32 | ~7.1 offload | — | Gemma 4 12B |
| RTX 2000 Ada Ada Lovelace · GDDR6 ECC · 2024 |
16 GB | 224 GB/s | 38 | 25 | ~6.2 offload | — | Gemma 4 12B |
Best 20–24 GB GPU for local LLMs
what to run on 20-24 GB →The sweet spot: a current dense 27B with vision and context. The RTX 3090 and 4090 are why "24 GB" is the number people quote.
| Device | Memory | Bandwidth | Models fit | Qwen3.5 9B tok/s | Qwen3.8 27B tok/s | Llama 3.3 70B Instruct tok/s | Best overall pick |
|---|---|---|---|---|---|---|---|
| GeForce RTX 4090 fastest in class Ada Lovelace · GDDR6X · 2022 |
24 GB | 1008 GB/s | 58 | 112 | 39 | ~1.7 offload | Qwen3.8 27B |
| GeForce RTX 3090 Ti Ampere · GDDR6X · 2022 |
24 GB | 1008 GB/s | 58 | 112 | 39 | ~1.7 offload | Qwen3.8 27B |
| Radeon RX 7900 XTX RDNA 3 · GDDR6 · 2022 |
24 GB | 960 GB/s | 58 | 107 | 37 | ~1.7 offload | Qwen3.8 27B |
| GeForce RTX 3090 Ampere · GDDR6X · 2020 |
24 GB | 936 GB/s | 58 | 104 | 36 | ~1.7 offload | Qwen3.8 27B |
| GeForce RTX 5090 Laptop Blackwell · GDDR7 · 2025 |
24 GB | 896 GB/s | 58 | 100 | 35 | ~1.7 offload | Qwen3.8 27B |
| RTX A5000 Ampere · GDDR6 ECC · 2021 |
24 GB | 768 GB/s | 58 | 86 | 30 | ~1.7 offload | Qwen3.8 27B |
| A10 Ampere · GDDR6 ECC · 2021 |
24 GB | 600 GB/s | 58 | 67 | 23 | ~1.6 offload | Qwen3.8 27B |
| Tesla P40 Pascal · GDDR5 · 2016 |
24 GB | 346 GB/s | 58 | 39 | 13 | ~1.5 offload | Qwen3.8 27B |
| L4 Ada Lovelace · GDDR6 ECC · 2023 |
24 GB | 300 GB/s | 58 | 33 | 12 | ~1.5 offload | Qwen3.8 27B |
| Radeon RX 7900 XT RDNA 3 · GDDR6 · 2022 |
20 GB | 800 GB/s | 45 | 89 | 31 | ~1.4 offload | Qwen3.8 27B |
| RTX 4000 Ada Ada Lovelace · GDDR6 ECC · 2023 |
20 GB | 360 GB/s | 45 | 40 | 14 | ~1.3 offload | Qwen3.8 27B |
Best 32 GB GPU for local LLMs
what to run on 32 GB →A 27B at Q6 or Q8, or the same model with a very long context.
| Device | Memory | Bandwidth | Models fit | Qwen3.5 9B tok/s | Qwen3.8 27B tok/s | Llama 3.3 70B Instruct tok/s | Best overall pick |
|---|---|---|---|---|---|---|---|
| GeForce RTX 5090 fastest in class Blackwell · GDDR7 · 2025 |
32 GB | 1792 GB/s | 58 | 200 | 69 | ~2.7 offload | Qwen3.8 27B |
| RTX 5000 Ada Ada Lovelace · GDDR6 ECC · 2023 |
32 GB | 576 GB/s | 58 | 64 | 22 | ~2.4 offload | Qwen3.8 27B |
| Radeon Pro W7800 RDNA 3 · GDDR6 ECC · 2023 |
32 GB | 576 GB/s | 58 | 64 | 22 | ~2.4 offload | Qwen3.8 27B |
Best 40–48 GB GPU for local LLMs
what to run on 40-48 GB →Workstation class. A 70B at Q4 fits tightly; a 30B near-lossless with 128K context fits easily.
| Device | Memory | Bandwidth | Models fit | Qwen3.5 9B tok/s | Qwen3.8 27B tok/s | Llama 3.3 70B Instruct tok/s | Best overall pick |
|---|---|---|---|---|---|---|---|
| RTX 6000 Ada Ada Lovelace · GDDR6 ECC · 2022 |
48 GB | 960 GB/s | 60 | 107 | 37 | 15 | Qwen3.8 27B |
| L40S Ada Lovelace · GDDR6 ECC · 2023 |
48 GB | 864 GB/s | 60 | 96 | 33 | 13 | Qwen3.8 27B |
| Radeon Pro W7900 RDNA 3 · GDDR6 ECC · 2023 |
48 GB | 864 GB/s | 60 | 96 | 33 | 13 | Qwen3.8 27B |
| RTX A6000 Ampere · GDDR6 ECC · 2020 |
48 GB | 768 GB/s | 60 | 86 | 30 | 12 | Qwen3.8 27B |
| A100 40 GB fastest in class Ampere · HBM2 · 2020 |
40 GB | 1555 GB/s | 58 | 173 | 60 | ~6.3 offload | Qwen3.8 27B |
Best 64–80 GB GPU for local LLMs
Data-centre cards: 70B at Q5–Q6, gpt-oss 120B whole.
| Device | Memory | Bandwidth | Models fit | Qwen3.5 9B tok/s | Qwen3.8 27B tok/s | Llama 3.3 70B Instruct tok/s | Best overall pick |
|---|---|---|---|---|---|---|---|
| H100 SXM fastest in class Hopper · HBM3 · 2022 |
80 GB | 3350 GB/s | 66 | 374 | 130 | 51 | Qwen3.8 27B |
| A100 80 GB Ampere · HBM2e · 2021 |
80 GB | 2039 GB/s | 66 | 227 | 79 | 31 | Qwen3.8 27B |
| H100 PCIe Hopper · HBM3 · 2022 |
80 GB | 2000 GB/s | 66 | 223 | 77 | 30 | Qwen3.8 27B |
| Instinct MI210 CDNA 2 · HBM2e · 2022 |
64 GB | 1638 GB/s | 61 | 183 | 63 | 25 | Qwen3.8 27B |
Best 128 GB and up GPU for local LLMs
The 100–300B mixture-of-experts models on one card.
| Device | Memory | Bandwidth | Models fit | Qwen3.5 9B tok/s | Qwen3.8 27B tok/s | Llama 3.3 70B Instruct tok/s | Best overall pick |
|---|---|---|---|---|---|---|---|
| Instinct MI300X fastest in class CDNA 3 · HBM3 · 2023 |
192 GB | 5300 GB/s | 70 | 591 | 205 | 81 | DeepSeek V4 Flash 284B-A13B |
| H200 SXM Hopper · HBM3e · 2024 |
141 GB | 4800 GB/s | 68 | 536 | 186 | 73 | Qwen3.8-Flash-Next 180B-A6B |
Unified memory: Apple Silicon and Strix Halo
One pool shared with the OS, of which the GPU may wire down roughly three quarters. Bandwidth, not capacity, is what separates an M4 Max from a Ryzen AI Max with the same 128 GB.
| Device | Memory | Bandwidth | Models fit | Qwen3.5 9B tok/s | Qwen3.8 27B tok/s | Llama 3.3 70B Instruct tok/s | Best overall pick |
|---|---|---|---|---|---|---|---|
| M3 Ultra · 512 GB M3 Ultra · LPDDR5X unified · 2025 |
512 GB | 800 GB/s | 74 | 82 | 29 | 11 | DeepSeek V4.1 Flash 552B-A16B |
| M3 Ultra · 256 GB M3 Ultra · LPDDR5X unified · 2025 |
256 GB | 800 GB/s | 70 | 82 | 29 | 11 | DeepSeek V4 Flash 284B-A13B |
| M2 Ultra · 192 GB M2 Ultra · LPDDR5 unified · 2023 |
192 GB | 800 GB/s | 68 | 82 | 29 | 11 | Qwen3.8-Flash-Next 180B-A6B |
| M1 Ultra · 128 GB M1 Ultra · LPDDR5 unified · 2022 |
128 GB | 800 GB/s | 66 | 82 | 29 | 11 | Qwen3.8 27B |
| M4 Max · 128 GB M4 Max · LPDDR5X unified · 2024 |
128 GB | 546 GB/s | 66 | 56 | 20 | 7.7 | Qwen3.8 27B |
| M3 Max · 128 GB M3 Max · LPDDR5 unified · 2023 |
128 GB | 400 GB/s | 66 | 41 | 14 | 5.6 | Qwen3.8 27B |
| Ryzen AI Max+ 395 · 128 GB Strix Halo · LPDDR5X unified · 2025 |
128 GB | 256 GB/s | 66 | 26 | 9.2 | 3.6 | Qwen3.8 27B |
| M2 Max · 96 GB M2 Max · LPDDR5 unified · 2023 |
96 GB | 400 GB/s | 66 | 41 | 14 | 5.6 | Qwen3.8 27B |
| M4 Max · 64 GB M4 Max · LPDDR5X unified · 2024 |
64 GB | 410 GB/s | 60 | 42 | 15 | 5.8 | Qwen3.8 27B |
| M1 Max · 64 GB M1 Max · LPDDR5 unified · 2021 |
64 GB | 400 GB/s | 60 | 41 | 14 | 5.6 | Qwen3.8 27B |
| M3 Max · 64 GB M3 Max · LPDDR5 unified · 2023 |
64 GB | 400 GB/s | 60 | 41 | 14 | 5.6 | Qwen3.8 27B |
| M4 Pro · 64 GB M4 Pro · LPDDR5X unified · 2024 |
64 GB | 273 GB/s | 60 | 28 | 9.8 | 3.8 | Qwen3.8 27B |
| Ryzen AI Max+ 395 · 64 GB Strix Halo · LPDDR5X unified · 2025 |
64 GB | 256 GB/s | 60 | 26 | 9.2 | 3.6 | Qwen3.8 27B |
| M4 Pro · 48 GB M4 Pro · LPDDR5X unified · 2024 |
48 GB | 273 GB/s | 58 | 28 | 9.8 | — | Qwen3.8 27B |
| M3 Pro · 36 GB M3 Pro · LPDDR5 unified · 2023 |
36 GB | 150 GB/s | 58 | 15 | 5.4 | — | Qwen3.8 27B |
| M1 Pro · 32 GB M1 Pro · LPDDR5 unified · 2021 |
32 GB | 200 GB/s | 58 | 21 | 7.1 | — | Qwen3.8 27B |
| M2 Pro · 32 GB M2 Pro · LPDDR5 unified · 2023 |
32 GB | 200 GB/s | 58 | 21 | 7.1 | — | Qwen3.8 27B |
| M4 · 32 GB M4 · LPDDR5X unified · 2024 |
32 GB | 120 GB/s | 58 | 12 | 4.3 | — | Qwen3.8 27B |
| M2 · 24 GB M2 · LPDDR5 unified · 2022 |
24 GB | 100 GB/s | 45 | 10 | 3.6 | — | Qwen3.8 27B |
| M3 · 24 GB M3 · LPDDR5 unified · 2023 |
24 GB | 100 GB/s | 45 | 10 | 3.6 | — | Qwen3.8 27B |
| M1 · 16 GB M1 · LPDDR4X unified · 2020 |
16 GB | 68 GB/s | 37 | 7.0 | — | — | Gemma 4 12B |
Mac or GPU?
A 64 GB M4 Max holds a 70B model at Q4 that no 24 GB card can; a 128 GB Mac holds 100B-class mixture-of-experts models that would otherwise need a data-centre card. What you give up is bandwidth: 410–546 GB/s against 1,000+ on a discrete flagship, so the same 27B runs at roughly a third of the speed. The rule of thumb this site's arithmetic supports: a Mac wins when the model you want does not fit your card at all, a discrete GPU wins when it does.
A used RTX 3090 remains the strongest argument in the table: 24 GB and 936 GB/s, the same class as a 4090 for everything a single user does. The Tesla P40 is the opposite warning — 24 GB, but at 346 GB/s a 27B runs at a third of the speed, and its Pascal cores lack the instructions the newer quantisations want.
How much GPU do I need?
| If you want to run | You need | Example |
|---|---|---|
| A small assistant, autocomplete, phone-class models Qwen3.5 4B, 4.66B |
8 GB card or a 16 GB Mac |
GeForce RTX 3070 Ti |
| A real everyday assistant with vision Qwen3.5 9B, 9.65B |
10 GB card or a 16 GB Mac |
GeForce RTX 3080 |
| A 12B generalist with headroom Gemma 4 12B, 12B |
11 GB card or a 16 GB Mac |
GeForce RTX 2080 Ti |
| gpt-oss 20B with its full 128K context gpt-oss 20B, 20.9B MoE |
16 GB card or a 24 GB Mac |
GeForce RTX 5080 |
| The current 27B generation: coding agents, vision, 262K context Qwen3.8 27B, 27.8B |
24 GB card or a 32 GB Mac |
GeForce RTX 4090 |
| A 70B dense model Llama 3.3 70B Instruct, 70.6B |
64 GB card or a 96 GB Mac |
Instinct MI210 |
| gpt-oss 120B whole gpt-oss 120B, 117B MoE |
80 GB card or a 128 GB Mac |
H100 SXM |
| The 300B-class MoEs GLM-5.3-Flash 320B-A18B, 321B MoE |
more than any single consumer card or a 512 GB Mac |
multi-GPU or unified memory |
"Needs" is the smallest discrete class that runs the model at the recommended quantisation with 15% headroom and an 8K context. Every model page has a what hardware do I need? companion that runs this for all 91 devices.
Whether to buy at all is a break-even question: a card is paid once, an API is paid per token. The AI article cost calculator prices the per-token side across 40 hosted models; hold it against the card you were about to buy.