Yes — with 21.7 GB to spare
Qwen3.5-27B at Q4_K_M fits your A100 40 GB entirely on the GPU at 8K context, at an estimated 60 tokens per second. There is room for its full 256K window.
Sized from Qwen/Qwen3.5-27B's config.json — not a hand-verified catalogue entry.
Read from config.json
Parameter count is safetensors index; the active count for mixture-of-experts models is computed from expert sizes and may differ from the model card by a few percent. Pipeline: image-text-to-text · 1,888,406 downloads.
GGUF conversions found
- unsloth/Qwen3.5-27B-GGUF 180,523 downloads
- bartowski/Qwen_Qwen3.5-27B-GGUF 163,457 downloads
- Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-GGUF 26,639 downloads
- wcn123/Qwen3.5-27B-WebNovel-Writer-zh-GGUF 17,792 downloads
- lmstudio-community/Qwen3.5-27B-GGUF 10,244 downloads
- unsloth/Qwen3.5-27B-MTP-GGUF 6,093 downloads
The llama.cpp command above uses the original repo id; point -hf at one of these if the original has no GGUF files.
The VRAM budget
Quantisation ladder
| Quant | Weights | Total @ 8K | Max context | Tok/s | Quality | Fit |
|---|---|---|---|---|---|---|
| Q8_0 | 27.5 GB | 28.6 GB | 164K | 34 | −0.1% ppl | Long context |
| Q6_K | 21.2 GB | 22.3 GB | 256K | 44 | −0.4% ppl | Long context |
| Q5_K_M | 18.3 GB | 19.4 GB | 256K | 51 | −0.8% ppl | Long context |
| Q4_K_M | 15.6 GB | 16.7 GB | 256K | 60 | −1.9% ppl | Recommended |
| Q3_K_M | 12.6 GB | 13.7 GB | 256K | 74 | −5.4% ppl | Long context |
| Q2_K | 10.8 GB | 11.9 GB | 256K | 87 | −15% ppl | Long context |
Quality is the published perplexity delta against f16 weights. Max context assumes an f16 KV cache; q8_0 roughly doubles it. Only 16 of its 64 blocks keep a per-token KV cache; the rest are linear-attention, Mamba or convolution blocks with a fixed-size state.
How to run it
$ llama-server \
-hf Qwen/Qwen3.5-27B:Q4_K_M \
-c 8192 -ngl 99
The engine underneath most of the others. Every knob is exposed. More on llama.cpp.