Best GPU for local LLMs · 2026

The best GPU for local LLMs, ranked by what it can actually run

91 GPUs, Apple Silicon chips and CPU configurations, ranked inside each memory class by what fits and how fast it runs. No prices — they move weekly and vary by country — and no benchmarks: the speeds are estimated from each card's published memory bandwidth, the same way for every device, so they are fair to compare and wrong in absolute terms by a similar factor.

1 · VRAM decides what runs

A model either fits in memory or it does not, and when it does not you lose an order of magnitude in speed to offload. Buy capacity first: 24 GB unlocks the current 27B generation, 16 GB the 20B MoEs, 12 GB a 12B with headroom, 8 GB a 9B tightly.

2 · Bandwidth decides how fast

Generation reads every active weight once per token, so tokens per second is bandwidth divided by model size. Within a class, that is the whole ranking — an RTX 3090's 936 GB/s runs a 27B at nearly the speed of a 4090's 1,008.

3 · Then the software

NVIDIA runs everything on day one. AMD and Intel run llama.cpp and Ollama well via Vulkan and ROCm, and vLLM on AMD only. Apple Silicon has MLX, which is faster than GGUF on the same Mac but lags on the newest releases.

Best 8 GB GPU for local LLMs

what to run on 8 GB →

Enough for a 4–9B assistant and the small MoEs. A 9B at Q4 is a tight fit; 7B and under run with headroom.

DeviceMemoryBandwidthModels fitQwen3.5 9B tok/sQwen3.8 27B tok/sLlama 3.3 70B Instruct tok/sBest overall pick
GeForce RTX 3070 Ti fastest in class
Ampere · GDDR6X · 2021
8 GB 608 GB/s 28 68~3.5 offload— Gemma 4 E4B
Arc A750
Alchemist · GDDR6 · 2022
8 GB 512 GB/s 28 57~3.5 offload— Gemma 4 E4B
GeForce RTX 5060
Blackwell · GDDR7 · 2025
8 GB 448 GB/s 28 50~3.5 offload— Gemma 4 E4B
GeForce RTX 3070
Ampere · GDDR6 · 2020
8 GB 448 GB/s 28 50~3.5 offload— Gemma 4 E4B
GeForce RTX 3060 Ti
Ampere · GDDR6 · 2020
8 GB 448 GB/s 28 50~3.5 offload— Gemma 4 E4B
GeForce RTX 4060 Ti
Ada Lovelace · GDDR6 · 2023
8 GB 288 GB/s 28 32~3.3 offload— Gemma 4 E4B
GeForce RTX 4060
Ada Lovelace · GDDR6 · 2023
8 GB 272 GB/s 28 30~3.3 offload— Gemma 4 E4B
GeForce RTX 4070 Laptop
Ada Lovelace · GDDR6 · 2023
8 GB 256 GB/s 28 29~3.3 offload— Gemma 4 E4B

Best 10–12 GB GPU for local LLMs

what to run on 10-12 GB →

The first class where a 12B runs with headroom. Used RTX 3060 12 GB cards made this the default budget class.

DeviceMemoryBandwidthModels fitQwen3.5 9B tok/sQwen3.8 27B tok/sLlama 3.3 70B Instruct tok/sBest overall pick
GeForce RTX 3080 Ti fastest in class
Ampere · GDDR6X · 2021
12 GB 912 GB/s 37 102~5.4 offload— Gemma 4 12B
GeForce RTX 3080 12 GB
Ampere · GDDR6X · 2022
12 GB 912 GB/s 37 102~5.4 offload— Gemma 4 12B
GeForce RTX 5070
Blackwell · GDDR7 · 2025
12 GB 672 GB/s 37 75~5.2 offload— Gemma 4 12B
GeForce RTX 4070 Ti
Ada Lovelace · GDDR6X · 2023
12 GB 504 GB/s 37 56~5.0 offload— Gemma 4 12B
GeForce RTX 4070 Super
Ada Lovelace · GDDR6X · 2024
12 GB 504 GB/s 37 56~5.0 offload— Gemma 4 12B
GeForce RTX 4070
Ada Lovelace · GDDR6X · 2023
12 GB 504 GB/s 37 56~5.0 offload— Gemma 4 12B
Arc B580
Battlemage · GDDR6 · 2024
12 GB 456 GB/s 37 51~4.9 offload— Gemma 4 12B
GeForce RTX 4080 Laptop
Ada Lovelace · GDDR6 · 2023
12 GB 432 GB/s 37 48~4.9 offload— Gemma 4 12B
Radeon RX 7700 XT
RDNA 3 · GDDR6 · 2023
12 GB 432 GB/s 37 48~4.9 offload— Gemma 4 12B
Radeon RX 6700 XT
RDNA 2 · GDDR6 · 2021
12 GB 384 GB/s 37 43~4.8 offload— Gemma 4 12B
GeForce RTX 3060 12 GB
Ampere · GDDR6 · 2021
12 GB 360 GB/s 37 40~4.7 offload— Gemma 4 12B
GeForce RTX 2060 12 GB
Turing · GDDR6 · 2021
12 GB 336 GB/s 37 37~4.6 offload— Gemma 4 12B
GeForce RTX 2080 Ti
Turing · GDDR6 · 2018
11 GB 616 GB/s 33 69~4.6 offload— Gemma 4 12B
GeForce GTX 1080 Ti
Pascal · GDDR5X · 2017
11 GB 484 GB/s 33 54~4.5 offload— Gemma 4 12B
GeForce RTX 3080
Ampere · GDDR6X · 2020
10 GB 760 GB/s 32 85~4.3 offload— Qwen3.5 9B
Arc B570
Battlemage · GDDR6 · 2025
10 GB 380 GB/s 32 42~4.0 offload— Qwen3.5 9B

Best 16 GB GPU for local LLMs

what to run on 16 GB →

gpt-oss 20B with its full 128K window, 14B dense models with room to spare, 24B tightly.

DeviceMemoryBandwidthModels fitQwen3.5 9B tok/sQwen3.8 27B tok/sLlama 3.3 70B Instruct tok/sBest overall pick
GeForce RTX 5080 fastest in class
Blackwell · GDDR7 · 2025
16 GB 960 GB/s 38 107~11 offload— Gemma 4 12B
GeForce RTX 5070 Ti
Blackwell · GDDR7 · 2025
16 GB 896 GB/s 38 100~11 offload— Gemma 4 12B
GeForce RTX 4080 Super
Ada Lovelace · GDDR6X · 2024
16 GB 736 GB/s 38 82~11 offload— Gemma 4 12B
GeForce RTX 4080
Ada Lovelace · GDDR6X · 2022
16 GB 717 GB/s 38 80~11 offload— Gemma 4 12B
GeForce RTX 4070 Ti Super
Ada Lovelace · GDDR6X · 2024
16 GB 672 GB/s 38 75~10 offload— Gemma 4 12B
Radeon RX 9070 XT
RDNA 4 · GDDR6 · 2025
16 GB 645 GB/s 38 72~10 offload— Gemma 4 12B
Radeon RX 9070
RDNA 4 · GDDR6 · 2025
16 GB 645 GB/s 38 72~10 offload— Gemma 4 12B
Radeon RX 7800 XT
RDNA 3 · GDDR6 · 2023
16 GB 624 GB/s 38 70~10 offload— Gemma 4 12B
GeForce RTX 4090 Laptop
Ada Lovelace · GDDR6 · 2023
16 GB 576 GB/s 38 64~9.8 offload— Gemma 4 12B
Radeon RX 7900 GRE
RDNA 3 · GDDR6 · 2023
16 GB 576 GB/s 38 64~9.8 offload— Gemma 4 12B
Arc A770 16 GB
Alchemist · GDDR6 · 2022
16 GB 560 GB/s 38 62~9.7 offload— Gemma 4 12B
Radeon RX 6900 XT
RDNA 2 · GDDR6 · 2020
16 GB 512 GB/s 38 57~9.3 offload— Gemma 4 12B
Radeon RX 6800 XT
RDNA 2 · GDDR6 · 2020
16 GB 512 GB/s 38 57~9.3 offload— Gemma 4 12B
GeForce RTX 5060 Ti 16 GB
Blackwell · GDDR7 · 2025
16 GB 448 GB/s 38 50~8.8 offload— Gemma 4 12B
RTX A4000
Ampere · GDDR6 ECC · 2021
16 GB 448 GB/s 38 50~8.8 offload— Gemma 4 12B
GeForce RTX 4060 Ti 16 GB
Ada Lovelace · GDDR6 · 2023
16 GB 288 GB/s 38 32~7.1 offload— Gemma 4 12B
Radeon RX 7600 XT
RDNA 3 · GDDR6 · 2024
16 GB 288 GB/s 38 32~7.1 offload— Gemma 4 12B
RTX 2000 Ada
Ada Lovelace · GDDR6 ECC · 2024
16 GB 224 GB/s 38 25~6.2 offload— Gemma 4 12B

Best 20–24 GB GPU for local LLMs

what to run on 20-24 GB →

The sweet spot: a current dense 27B with vision and context. The RTX 3090 and 4090 are why "24 GB" is the number people quote.

DeviceMemoryBandwidthModels fitQwen3.5 9B tok/sQwen3.8 27B tok/sLlama 3.3 70B Instruct tok/sBest overall pick
GeForce RTX 4090 fastest in class
Ada Lovelace · GDDR6X · 2022
24 GB 1008 GB/s 58 11239~1.7 offload Qwen3.8 27B
GeForce RTX 3090 Ti
Ampere · GDDR6X · 2022
24 GB 1008 GB/s 58 11239~1.7 offload Qwen3.8 27B
Radeon RX 7900 XTX
RDNA 3 · GDDR6 · 2022
24 GB 960 GB/s 58 10737~1.7 offload Qwen3.8 27B
GeForce RTX 3090
Ampere · GDDR6X · 2020
24 GB 936 GB/s 58 10436~1.7 offload Qwen3.8 27B
GeForce RTX 5090 Laptop
Blackwell · GDDR7 · 2025
24 GB 896 GB/s 58 10035~1.7 offload Qwen3.8 27B
RTX A5000
Ampere · GDDR6 ECC · 2021
24 GB 768 GB/s 58 8630~1.7 offload Qwen3.8 27B
A10
Ampere · GDDR6 ECC · 2021
24 GB 600 GB/s 58 6723~1.6 offload Qwen3.8 27B
Tesla P40
Pascal · GDDR5 · 2016
24 GB 346 GB/s 58 3913~1.5 offload Qwen3.8 27B
L4
Ada Lovelace · GDDR6 ECC · 2023
24 GB 300 GB/s 58 3312~1.5 offload Qwen3.8 27B
Radeon RX 7900 XT
RDNA 3 · GDDR6 · 2022
20 GB 800 GB/s 45 8931~1.4 offload Qwen3.8 27B
RTX 4000 Ada
Ada Lovelace · GDDR6 ECC · 2023
20 GB 360 GB/s 45 4014~1.3 offload Qwen3.8 27B

Best 32 GB GPU for local LLMs

what to run on 32 GB →

A 27B at Q6 or Q8, or the same model with a very long context.

DeviceMemoryBandwidthModels fitQwen3.5 9B tok/sQwen3.8 27B tok/sLlama 3.3 70B Instruct tok/sBest overall pick
GeForce RTX 5090 fastest in class
Blackwell · GDDR7 · 2025
32 GB 1792 GB/s 58 20069~2.7 offload Qwen3.8 27B
RTX 5000 Ada
Ada Lovelace · GDDR6 ECC · 2023
32 GB 576 GB/s 58 6422~2.4 offload Qwen3.8 27B
Radeon Pro W7800
RDNA 3 · GDDR6 ECC · 2023
32 GB 576 GB/s 58 6422~2.4 offload Qwen3.8 27B

Best 40–48 GB GPU for local LLMs

what to run on 40-48 GB →

Workstation class. A 70B at Q4 fits tightly; a 30B near-lossless with 128K context fits easily.

DeviceMemoryBandwidthModels fitQwen3.5 9B tok/sQwen3.8 27B tok/sLlama 3.3 70B Instruct tok/sBest overall pick
RTX 6000 Ada
Ada Lovelace · GDDR6 ECC · 2022
48 GB 960 GB/s 60 1073715 Qwen3.8 27B
L40S
Ada Lovelace · GDDR6 ECC · 2023
48 GB 864 GB/s 60 963313 Qwen3.8 27B
Radeon Pro W7900
RDNA 3 · GDDR6 ECC · 2023
48 GB 864 GB/s 60 963313 Qwen3.8 27B
RTX A6000
Ampere · GDDR6 ECC · 2020
48 GB 768 GB/s 60 863012 Qwen3.8 27B
A100 40 GB fastest in class
Ampere · HBM2 · 2020
40 GB 1555 GB/s 58 17360~6.3 offload Qwen3.8 27B

Best 64–80 GB GPU for local LLMs

Data-centre cards: 70B at Q5–Q6, gpt-oss 120B whole.

DeviceMemoryBandwidthModels fitQwen3.5 9B tok/sQwen3.8 27B tok/sLlama 3.3 70B Instruct tok/sBest overall pick
H100 SXM fastest in class
Hopper · HBM3 · 2022
80 GB 3350 GB/s 66 37413051 Qwen3.8 27B
A100 80 GB
Ampere · HBM2e · 2021
80 GB 2039 GB/s 66 2277931 Qwen3.8 27B
H100 PCIe
Hopper · HBM3 · 2022
80 GB 2000 GB/s 66 2237730 Qwen3.8 27B
Instinct MI210
CDNA 2 · HBM2e · 2022
64 GB 1638 GB/s 61 1836325 Qwen3.8 27B

Best 128 GB and up GPU for local LLMs

The 100–300B mixture-of-experts models on one card.

DeviceMemoryBandwidthModels fitQwen3.5 9B tok/sQwen3.8 27B tok/sLlama 3.3 70B Instruct tok/sBest overall pick
Instinct MI300X fastest in class
CDNA 3 · HBM3 · 2023
192 GB 5300 GB/s 70 59120581 DeepSeek V4 Flash 284B-A13B
H200 SXM
Hopper · HBM3e · 2024
141 GB 4800 GB/s 68 53618673 Qwen3.8-Flash-Next 180B-A6B

Unified memory: Apple Silicon and Strix Halo

One pool shared with the OS, of which the GPU may wire down roughly three quarters. Bandwidth, not capacity, is what separates an M4 Max from a Ryzen AI Max with the same 128 GB.

DeviceMemoryBandwidthModels fitQwen3.5 9B tok/sQwen3.8 27B tok/sLlama 3.3 70B Instruct tok/sBest overall pick
M3 Ultra · 512 GB
M3 Ultra · LPDDR5X unified · 2025
512 GB 800 GB/s 74 822911 DeepSeek V4.1 Flash 552B-A16B
M3 Ultra · 256 GB
M3 Ultra · LPDDR5X unified · 2025
256 GB 800 GB/s 70 822911 DeepSeek V4 Flash 284B-A13B
M2 Ultra · 192 GB
M2 Ultra · LPDDR5 unified · 2023
192 GB 800 GB/s 68 822911 Qwen3.8-Flash-Next 180B-A6B
M1 Ultra · 128 GB
M1 Ultra · LPDDR5 unified · 2022
128 GB 800 GB/s 66 822911 Qwen3.8 27B
M4 Max · 128 GB
M4 Max · LPDDR5X unified · 2024
128 GB 546 GB/s 66 56207.7 Qwen3.8 27B
M3 Max · 128 GB
M3 Max · LPDDR5 unified · 2023
128 GB 400 GB/s 66 41145.6 Qwen3.8 27B
Ryzen AI Max+ 395 · 128 GB
Strix Halo · LPDDR5X unified · 2025
128 GB 256 GB/s 66 269.23.6 Qwen3.8 27B
M2 Max · 96 GB
M2 Max · LPDDR5 unified · 2023
96 GB 400 GB/s 66 41145.6 Qwen3.8 27B
M4 Max · 64 GB
M4 Max · LPDDR5X unified · 2024
64 GB 410 GB/s 60 42155.8 Qwen3.8 27B
M1 Max · 64 GB
M1 Max · LPDDR5 unified · 2021
64 GB 400 GB/s 60 41145.6 Qwen3.8 27B
M3 Max · 64 GB
M3 Max · LPDDR5 unified · 2023
64 GB 400 GB/s 60 41145.6 Qwen3.8 27B
M4 Pro · 64 GB
M4 Pro · LPDDR5X unified · 2024
64 GB 273 GB/s 60 289.83.8 Qwen3.8 27B
Ryzen AI Max+ 395 · 64 GB
Strix Halo · LPDDR5X unified · 2025
64 GB 256 GB/s 60 269.23.6 Qwen3.8 27B
M4 Pro · 48 GB
M4 Pro · LPDDR5X unified · 2024
48 GB 273 GB/s 58 289.8— Qwen3.8 27B
M3 Pro · 36 GB
M3 Pro · LPDDR5 unified · 2023
36 GB 150 GB/s 58 155.4— Qwen3.8 27B
M1 Pro · 32 GB
M1 Pro · LPDDR5 unified · 2021
32 GB 200 GB/s 58 217.1— Qwen3.8 27B
M2 Pro · 32 GB
M2 Pro · LPDDR5 unified · 2023
32 GB 200 GB/s 58 217.1— Qwen3.8 27B
M4 · 32 GB
M4 · LPDDR5X unified · 2024
32 GB 120 GB/s 58 124.3— Qwen3.8 27B
M2 · 24 GB
M2 · LPDDR5 unified · 2022
24 GB 100 GB/s 45 103.6— Qwen3.8 27B
M3 · 24 GB
M3 · LPDDR5 unified · 2023
24 GB 100 GB/s 45 103.6— Qwen3.8 27B
M1 · 16 GB
M1 · LPDDR4X unified · 2020
16 GB 68 GB/s 37 7.0—— Gemma 4 12B

Mac or GPU?

A 64 GB M4 Max holds a 70B model at Q4 that no 24 GB card can; a 128 GB Mac holds 100B-class mixture-of-experts models that would otherwise need a data-centre card. What you give up is bandwidth: 410–546 GB/s against 1,000+ on a discrete flagship, so the same 27B runs at roughly a third of the speed. The rule of thumb this site's arithmetic supports: a Mac wins when the model you want does not fit your card at all, a discrete GPU wins when it does.

A used RTX 3090 remains the strongest argument in the table: 24 GB and 936 GB/s, the same class as a 4090 for everything a single user does. The Tesla P40 is the opposite warning — 24 GB, but at 346 GB/s a 27B runs at a third of the speed, and its Pascal cores lack the instructions the newer quantisations want.

How much GPU do I need?

If you want to runYou needExample
A small assistant, autocomplete, phone-class models
Qwen3.5 4B, 4.66B
8 GB card
or a 16 GB Mac
GeForce RTX 3070 Ti
A real everyday assistant with vision
Qwen3.5 9B, 9.65B
10 GB card
or a 16 GB Mac
GeForce RTX 3080
A 12B generalist with headroom 11 GB card
or a 16 GB Mac
GeForce RTX 2080 Ti
gpt-oss 20B with its full 128K context
gpt-oss 20B, 20.9B MoE
16 GB card
or a 24 GB Mac
GeForce RTX 5080
The current 27B generation: coding agents, vision, 262K context 24 GB card
or a 32 GB Mac
GeForce RTX 4090
A 70B dense model 64 GB card
or a 96 GB Mac
Instinct MI210
gpt-oss 120B whole
gpt-oss 120B, 117B MoE
80 GB card
or a 128 GB Mac
H100 SXM
The 300B-class MoEs more than any single consumer card
or a 512 GB Mac
multi-GPU or unified memory

"Needs" is the smallest discrete class that runs the model at the recommended quantisation with 15% headroom and an 8K context. Every model page has a what hardware do I need? companion that runs this for all 91 devices.


Whether to buy at all is a break-even question: a card is paid once, an API is paid per token. The AI article cost calculator prices the per-token side across 40 hosted models; hold it against the card you were about to buy.