Guides
Guides, picks and calculators
The tool answers "can my hardware run it?". These pages answer the questions around it — what to buy, what to run, what the file names mean — from the same catalogues and the same arithmetic, so they update when the catalogue does.
Pick hardware, pick a model, pick a runtime. Self-hosting on a GPU, a Mac or a CPU, with the commands.
ExplainerGGUF and quantization explainedWhat Q4_K_M, Q8_0 and the rest mean, how much each saves, and which file to download.
PicksBest local LLM by VRAMOne pick per job for every memory class from 8 GB to 128 GB.
PicksBest local LLM for codingThe coding shortlist, ranked per VRAM class, with max context for agents.
HardwareBest GPU for local LLMsEvery card and Mac ranked by what it can run and how fast.
ToolLLM VRAM calculatorAny model against any device: weights, KV cache, overhead, verdict.
ToolKV cache calculatorBytes per token and cache size at 4K, 8K, 32K and 128K for every model.
CatalogueOllama models that fit your GPUEvery library tag, its download size and the card it needs.
RuntimeLM Studio hardware requirementsThe app's minimums, then the memory each model takes, with the lms commands.