Hardware requirements

What hardware do I need to run gpt-oss 120B?

117B MoE · 5.1B active 128K context Apache 2.0

At MXFP4 the weights are 60.5 GB; with an 8K KV cache and runtime overhead that is 61.4 GB, so it wants a 80 GB card to run with headroom — in practice 80 GB VRAM: H100 SXM and 2 others. Every figure below is computed from the model’s published shape against each device’s usable memory; tokens per second are estimates. The method.

Minimum practical 64 GB VRAM

Instinct MI210 — MXFP4, 61.4 GB of 62.4 GB, ~144 tok/s, up to 36K context.

the smallest card that holds it at all, at Q3_K_M — the lowest quantisation worth running
Recommended 80 GB VRAM

H100 SXM — MXFP4, 61.4 GB of 78.4 GB, ~295 tok/s, up to 128K context.

runs the recommended quantisation with headroom at 8K context
High quality Nothing in the catalogue

No single device we list clears this bar — this is a multi-GPU or server-class model.

Long context 80 GB VRAM

H100 SXM — MXFP4, 65.6 GB of 78.4 GB, ~295 tok/s at 128K.

128K tokens (or the model’s maximum) at the recommended quantisation
Apple Silicon 128 GB unified

M1 Ultra · 128 GB — MXFP4, 61.4 GB of 96.0 GB, ~79 tok/s, up to 128K context.

smallest Mac that runs the recommended quantisation with headroom
CPU only 256 GB RAM

CPU only · DDR5 8-channel server — MXFP4, 61.4 GB of 204.8 GB, ~22 tok/s, up to 128K context.

no GPU — weights stream from system RAM

How much VRAM at each quantisation

QuantWeights+KV 8K+KV 32K+KV 128KTotal @ 8KNeedsQuality
MXFP460.5 GB0.29 GB1.13 GB4.5 GB 61.4 GB80 GB cardReference

Totals add 0.6 GB runtime overhead and an f16 KV cache; "needs" is the smallest discrete card class whose usable memory (after display and driver reserve) leaves 15% headroom. A q8_0 cache roughly halves the KV columns.

Every device, checked

System RAM for offload:

NVIDIA · GeForce

DeviceMemoryBest quant with headroomFits at allMax context @ MXFP4Tok/s est.128K
GeForce RTX 5090 32 GB offload ~10 no Check
GeForce RTX 5080 16 GB offload ~6.7 no Check
GeForce RTX 5070 Ti 16 GB offload ~6.7 no Check
GeForce RTX 5070 12 GB offload ~6.2 no Check
GeForce RTX 5060 Ti 16 GB 16 GB offload ~6.6 no Check
GeForce RTX 5060 8 GB offload ~5.8 no Check
GeForce RTX 4090 24 GB offload ~7.9 no Check
GeForce RTX 4080 Super 16 GB offload ~6.7 no Check
GeForce RTX 4080 16 GB offload ~6.6 no Check
GeForce RTX 4070 Ti Super 16 GB offload ~6.6 no Check
GeForce RTX 4070 Ti 12 GB offload ~6.2 no Check
GeForce RTX 4070 Super 12 GB offload ~6.2 no Check
GeForce RTX 4070 12 GB offload ~6.2 no Check
GeForce RTX 4060 Ti 16 GB 16 GB offload ~6.4 no Check
GeForce RTX 4060 Ti 8 GB offload ~5.7 no Check
GeForce RTX 4060 8 GB offload ~5.7 no Check
GeForce RTX 3090 Ti 24 GB offload ~7.9 no Check
GeForce RTX 3090 24 GB offload ~7.9 no Check
GeForce RTX 3080 Ti 12 GB offload ~6.2 no Check
GeForce RTX 3080 12 GB 12 GB offload ~6.2 no Check
GeForce RTX 3080 10 GB offload ~6.0 no Check
GeForce RTX 3070 Ti 8 GB offload ~5.8 no Check
GeForce RTX 3070 8 GB offload ~5.8 no Check
GeForce RTX 3060 Ti 8 GB offload ~5.8 no Check
GeForce RTX 3060 12 GB 12 GB offload ~6.1 no Check
GeForce RTX 2080 Ti 11 GB offload ~6.1 no Check
GeForce RTX 2060 12 GB 12 GB offload ~6.1 no Check
GeForce GTX 1080 Ti 11 GB offload ~6.1 no Check

NVIDIA · GeForce laptop

DeviceMemoryBest quant with headroomFits at allMax context @ MXFP4Tok/s est.128K
GeForce RTX 5090 Laptop 24 GB offload ~7.9 no Check
GeForce RTX 4090 Laptop 16 GB offload ~6.6 no Check
GeForce RTX 4080 Laptop 12 GB offload ~6.1 no Check
GeForce RTX 4070 Laptop 8 GB offload ~5.7 no Check

NVIDIA · Workstation

DeviceMemoryBest quant with headroomFits at allMax context @ MXFP4Tok/s est.128K
RTX 6000 Ada 48 GB offload ~18 no Check
RTX 5000 Ada 32 GB offload ~9.4 no Check
RTX 4000 Ada 20 GB offload ~7.0 no Check
RTX 2000 Ada 16 GB offload ~6.3 no Check
RTX A6000 48 GB offload ~17 no Check
RTX A5000 24 GB offload ~7.9 no Check
RTX A4000 16 GB offload ~6.6 no Check

NVIDIA · Data centre

DeviceMemoryBest quant with headroomFits at allMax context @ MXFP4Tok/s est.128K
L40S 48 GB offload ~18 no Check
L4 24 GB offload ~7.4 no Check
A10 24 GB offload ~7.8 no Check
A100 40 GB 40 GB offload ~13 no Check
A100 80 GB 80 GB MXFP4 MXFP4 128K 180 yes Check
H100 PCIe 80 GB MXFP4 MXFP4 128K 176 yes Check
H100 SXM 80 GB MXFP4 MXFP4 128K 295 yes Check
H200 SXM 141 GB MXFP4 MXFP4 128K 423 yes Check
Tesla P40 24 GB offload ~7.5 no Check

AMD · Radeon

DeviceMemoryBest quant with headroomFits at allMax context @ MXFP4Tok/s est.128K
Radeon RX 9070 XT 16 GB offload ~6.6 no Check
Radeon RX 9070 16 GB offload ~6.6 no Check
Radeon RX 7900 XTX 24 GB offload ~7.9 no Check
Radeon RX 7900 XT 20 GB offload ~7.2 no Check
Radeon RX 7900 GRE 16 GB offload ~6.6 no Check
Radeon RX 7800 XT 16 GB offload ~6.6 no Check
Radeon RX 7700 XT 12 GB offload ~6.1 no Check
Radeon RX 7600 XT 16 GB offload ~6.4 no Check
Radeon RX 6900 XT 16 GB offload ~6.6 no Check
Radeon RX 6800 XT 16 GB offload ~6.6 no Check
Radeon RX 6700 XT 12 GB offload ~6.1 no Check

AMD · Radeon Pro

DeviceMemoryBest quant with headroomFits at allMax context @ MXFP4Tok/s est.128K
Radeon Pro W7900 48 GB offload ~18 no Check
Radeon Pro W7800 32 GB offload ~9.4 no Check

AMD · Instinct

DeviceMemoryBest quant with headroomFits at allMax context @ MXFP4Tok/s est.128K
Instinct MI210 64 GB MXFP4 36K 144 no Check
Instinct MI300X 192 GB MXFP4 MXFP4 128K 467 yes Check

AMD · Ryzen AI Max

DeviceMemoryBest quant with headroomFits at allMax context @ MXFP4Tok/s est.128K
Ryzen AI Max+ 395 · 128 GB 128 GB MXFP4 MXFP4 128K 25 yes Check
Ryzen AI Max+ 395 · 64 GB 64 GB no no Check

Apple · Apple Silicon

DeviceMemoryBest quant with headroomFits at allMax context @ MXFP4Tok/s est.128K
M1 · 16 GB 16 GB no no Check
M1 Pro · 32 GB 32 GB no no Check
M1 Max · 64 GB 64 GB no no Check
M1 Ultra · 128 GB 128 GB MXFP4 MXFP4 128K 79 yes Check
M2 · 24 GB 24 GB no no Check
M2 Pro · 32 GB 32 GB no no Check
M2 Max · 96 GB 96 GB MXFP4 128K 39 tight Check
M2 Ultra · 192 GB 192 GB MXFP4 MXFP4 128K 79 yes Check
M3 · 24 GB 24 GB no no Check
M3 Pro · 36 GB 36 GB no no Check
M3 Max · 64 GB 64 GB no no Check
M3 Max · 128 GB 128 GB MXFP4 MXFP4 128K 39 yes Check
M3 Ultra · 256 GB 256 GB MXFP4 MXFP4 128K 79 yes Check
M3 Ultra · 512 GB 512 GB MXFP4 MXFP4 128K 79 yes Check
M4 · 32 GB 32 GB no no Check
M4 Pro · 48 GB 48 GB no no Check
M4 Pro · 64 GB 64 GB no no Check
M4 Max · 64 GB 64 GB no no Check
M4 Max · 128 GB 128 GB MXFP4 MXFP4 128K 54 yes Check

Intel · Arc

DeviceMemoryBest quant with headroomFits at allMax context @ MXFP4Tok/s est.128K
Arc B580 12 GB offload ~6.1 no Check
Arc B570 10 GB offload ~5.9 no Check
Arc A770 16 GB 16 GB offload ~6.6 no Check
Arc A750 8 GB offload ~5.8 no Check

CPU · CPU + system RAM

DeviceMemoryBest quant with headroomFits at allMax context @ MXFP4Tok/s est.128K
CPU only · DDR4 dual-channel 32 GB no no Check
CPU only · DDR5 dual-channel 64 GB no no Check
CPU only · DDR5 8-channel server 256 GB MXFP4 MXFP4 128K 22 yes Check