Hardware requirements

What hardware do I need to run gpt-oss 20B?

20.9B MoE · 3.6B active 128K context Apache 2.0

At MXFP4 the weights are 10.8 GB; with an 8K KV cache and runtime overhead that is 11.6 GB, so it wants a 16 GB card to run with headroom — in practice 16 GB VRAM: GeForce RTX 5080 and 17 others. Every figure below is computed from the model’s published shape against each device’s usable memory; tokens per second are estimates. The method.

Minimum practical 16 GB VRAM

GeForce RTX 5080 — MXFP4, 11.6 GB of 14.4 GB, ~120 tok/s, up to 126K context.

the smallest card that holds it at all, at Q3_K_M — the lowest quantisation worth running
Recommended 16 GB VRAM

GeForce RTX 5080 — MXFP4, 11.6 GB of 14.4 GB, ~120 tok/s, up to 126K context.

runs the recommended quantisation with headroom at 8K context
High quality Nothing in the catalogue

No single device we list clears this bar — this is a multi-GPU or server-class model.

Long context 20 GB VRAM

Radeon RX 7900 XT — MXFP4, 14.4 GB of 18.4 GB, ~100 tok/s at 128K.

128K tokens (or the model’s maximum) at the recommended quantisation
Apple Silicon 24 GB unified

M2 · 24 GB — MXFP4, 11.6 GB of 18.0 GB, ~14 tok/s, up to 128K context.

smallest Mac that runs the recommended quantisation with headroom
CPU only 32 GB RAM

CPU only · DDR4 dual-channel — MXFP4, 11.6 GB of 25.6 GB, ~5.1 tok/s, up to 128K context.

no GPU — weights stream from system RAM

How much VRAM at each quantisation

QuantWeights+KV 8K+KV 32K+KV 128KTotal @ 8KNeedsQuality
MXFP410.8 GB0.19 GB0.75 GB3.0 GB 11.6 GB16 GB cardReference

Totals add 0.6 GB runtime overhead and an f16 KV cache; "needs" is the smallest discrete card class whose usable memory (after display and driver reserve) leaves 15% headroom. A q8_0 cache roughly halves the KV columns.

Every device, checked

System RAM for offload:

NVIDIA · GeForce

DeviceMemoryBest quant with headroomFits at allMax context @ MXFP4Tok/s est.128K
GeForce RTX 5090 32 GB MXFP4 MXFP4 128K 224 yes Check
GeForce RTX 5080 16 GB MXFP4 MXFP4 126K 120 no Check
GeForce RTX 5070 Ti 16 GB MXFP4 MXFP4 126K 112 no Check
GeForce RTX 5070 12 GB offload ~43 no Check
GeForce RTX 5060 Ti 16 GB 16 GB MXFP4 MXFP4 126K 56 no Check
GeForce RTX 5060 8 GB offload ~15 no Check
GeForce RTX 4090 24 GB MXFP4 MXFP4 128K 126 yes Check
GeForce RTX 4080 Super 16 GB MXFP4 MXFP4 126K 92 no Check
GeForce RTX 4080 16 GB MXFP4 MXFP4 126K 90 no Check
GeForce RTX 4070 Ti Super 16 GB MXFP4 MXFP4 126K 84 no Check
GeForce RTX 4070 Ti 12 GB offload ~37 no Check
GeForce RTX 4070 Super 12 GB offload ~37 no Check
GeForce RTX 4070 12 GB offload ~37 no Check
GeForce RTX 4060 Ti 16 GB 16 GB MXFP4 MXFP4 126K 36 no Check
GeForce RTX 4060 Ti 8 GB offload ~14 no Check
GeForce RTX 4060 8 GB offload ~14 no Check
GeForce RTX 3090 Ti 24 GB MXFP4 MXFP4 128K 126 yes Check
GeForce RTX 3090 24 GB MXFP4 MXFP4 128K 117 yes Check
GeForce RTX 3080 Ti 12 GB offload ~49 no Check
GeForce RTX 3080 12 GB 12 GB offload ~49 no Check
GeForce RTX 3080 10 GB offload ~24 no Check
GeForce RTX 3070 Ti 8 GB offload ~16 no Check
GeForce RTX 3070 8 GB offload ~15 no Check
GeForce RTX 3060 Ti 8 GB offload ~15 no Check
GeForce RTX 3060 12 GB 12 GB offload ~31 no Check
GeForce RTX 2080 Ti 11 GB offload ~29 no Check
GeForce RTX 2060 12 GB 12 GB offload ~29 no Check
GeForce GTX 1080 Ti 11 GB offload ~27 no Check

NVIDIA · GeForce laptop

DeviceMemoryBest quant with headroomFits at allMax context @ MXFP4Tok/s est.128K
GeForce RTX 5090 Laptop 24 GB MXFP4 MXFP4 128K 112 yes Check
GeForce RTX 4090 Laptop 16 GB MXFP4 MXFP4 126K 72 no Check
GeForce RTX 4080 Laptop 12 GB offload ~34 no Check
GeForce RTX 4070 Laptop 8 GB offload ~13 no Check

NVIDIA · Workstation

DeviceMemoryBest quant with headroomFits at allMax context @ MXFP4Tok/s est.128K
RTX 6000 Ada 48 GB MXFP4 MXFP4 128K 120 yes Check
RTX 5000 Ada 32 GB MXFP4 MXFP4 128K 72 yes Check
RTX 4000 Ada 20 GB MXFP4 MXFP4 128K 45 yes Check
RTX 2000 Ada 16 GB MXFP4 MXFP4 126K 28 no Check
RTX A6000 48 GB MXFP4 MXFP4 128K 96 yes Check
RTX A5000 24 GB MXFP4 MXFP4 128K 96 yes Check
RTX A4000 16 GB MXFP4 MXFP4 126K 56 no Check

NVIDIA · Data centre

DeviceMemoryBest quant with headroomFits at allMax context @ MXFP4Tok/s est.128K
L40S 48 GB MXFP4 MXFP4 128K 108 yes Check
L4 24 GB MXFP4 MXFP4 128K 37 yes Check
A10 24 GB MXFP4 MXFP4 128K 75 yes Check
A100 40 GB 40 GB MXFP4 MXFP4 128K 194 yes Check
A100 80 GB 80 GB MXFP4 MXFP4 128K 255 yes Check
H100 PCIe 80 GB MXFP4 MXFP4 128K 250 yes Check
H100 SXM 80 GB MXFP4 MXFP4 128K 418 yes Check
H200 SXM 141 GB MXFP4 MXFP4 128K 599 yes Check
Tesla P40 24 GB MXFP4 MXFP4 128K 43 yes Check

AMD · Radeon

DeviceMemoryBest quant with headroomFits at allMax context @ MXFP4Tok/s est.128K
Radeon RX 9070 XT 16 GB MXFP4 MXFP4 126K 81 no Check
Radeon RX 9070 16 GB MXFP4 MXFP4 126K 81 no Check
Radeon RX 7900 XTX 24 GB MXFP4 MXFP4 128K 120 yes Check
Radeon RX 7900 XT 20 GB MXFP4 MXFP4 128K 100 yes Check
Radeon RX 7900 GRE 16 GB MXFP4 MXFP4 126K 72 no Check
Radeon RX 7800 XT 16 GB MXFP4 MXFP4 126K 78 no Check
Radeon RX 7700 XT 12 GB offload ~34 no Check
Radeon RX 7600 XT 16 GB MXFP4 MXFP4 126K 36 no Check
Radeon RX 6900 XT 16 GB MXFP4 MXFP4 126K 64 no Check
Radeon RX 6800 XT 16 GB MXFP4 MXFP4 126K 64 no Check
Radeon RX 6700 XT 12 GB offload ~32 no Check

AMD · Radeon Pro

DeviceMemoryBest quant with headroomFits at allMax context @ MXFP4Tok/s est.128K
Radeon Pro W7900 48 GB MXFP4 MXFP4 128K 108 yes Check
Radeon Pro W7800 32 GB MXFP4 MXFP4 128K 72 yes Check

AMD · Instinct

DeviceMemoryBest quant with headroomFits at allMax context @ MXFP4Tok/s est.128K
Instinct MI210 64 GB MXFP4 MXFP4 128K 204 yes Check
Instinct MI300X 192 GB MXFP4 MXFP4 128K 662 yes Check

AMD · Ryzen AI Max

DeviceMemoryBest quant with headroomFits at allMax context @ MXFP4Tok/s est.128K
Ryzen AI Max+ 395 · 128 GB 128 GB MXFP4 MXFP4 128K 36 yes Check
Ryzen AI Max+ 395 · 64 GB 64 GB MXFP4 MXFP4 128K 36 yes Check

Apple · Apple Silicon

DeviceMemoryBest quant with headroomFits at allMax context @ MXFP4Tok/s est.128K
M1 · 16 GB 16 GB no no Check
M1 Pro · 32 GB 32 GB MXFP4 MXFP4 128K 28 yes Check
M1 Max · 64 GB 64 GB MXFP4 MXFP4 128K 56 yes Check
M1 Ultra · 128 GB 128 GB MXFP4 MXFP4 128K 112 yes Check
M2 · 24 GB 24 GB MXFP4 MXFP4 128K 14 yes Check
M2 Pro · 32 GB 32 GB MXFP4 MXFP4 128K 28 yes Check
M2 Max · 96 GB 96 GB MXFP4 MXFP4 128K 56 yes Check
M2 Ultra · 192 GB 192 GB MXFP4 MXFP4 128K 112 yes Check
M3 · 24 GB 24 GB MXFP4 MXFP4 128K 14 yes Check
M3 Pro · 36 GB 36 GB MXFP4 MXFP4 128K 21 yes Check
M3 Max · 64 GB 64 GB MXFP4 MXFP4 128K 56 yes Check
M3 Max · 128 GB 128 GB MXFP4 MXFP4 128K 56 yes Check
M3 Ultra · 256 GB 256 GB MXFP4 MXFP4 128K 112 yes Check
M3 Ultra · 512 GB 512 GB MXFP4 MXFP4 128K 112 yes Check
M4 · 32 GB 32 GB MXFP4 MXFP4 128K 17 yes Check
M4 Pro · 48 GB 48 GB MXFP4 MXFP4 128K 38 yes Check
M4 Pro · 64 GB 64 GB MXFP4 MXFP4 128K 38 yes Check
M4 Max · 64 GB 64 GB MXFP4 MXFP4 128K 57 yes Check
M4 Max · 128 GB 128 GB MXFP4 MXFP4 128K 76 yes Check

Intel · Arc

DeviceMemoryBest quant with headroomFits at allMax context @ MXFP4Tok/s est.128K
Arc B580 12 GB offload ~35 no Check
Arc B570 10 GB offload ~20 no Check
Arc A770 16 GB 16 GB MXFP4 MXFP4 126K 70 no Check
Arc A750 8 GB offload ~15 no Check

CPU · CPU + system RAM

DeviceMemoryBest quant with headroomFits at allMax context @ MXFP4Tok/s est.128K
CPU only · DDR4 dual-channel 32 GB MXFP4 MXFP4 128K 5.1 yes Check
CPU only · DDR5 dual-channel 64 GB MXFP4 MXFP4 128K 9.0 yes Check
CPU only · DDR5 8-channel server 256 GB MXFP4 MXFP4 128K 31 yes Check