LLM Hardware Recommender

Free tool. Pick a model (Llama 3.1 / Qwen 2.5 / Mistral / DeepSeek), set your minimum tokens/sec, get the cheapest Mac M-series or NVIDIA / AMD GPU that runs it smoothly. 80+ measured benchmarks.

Tell us which open-weight model you want to run smoothly. We'll search every Mac M-series and every NVIDIA / AMD GPU in our catalog, look up real benchmarks, and tell you the cheapest setup that hits your speed target.

💰 Cheapest that meets your bar
RTX 6000 Ada
$7,600
18.5 tok/s · Q4_K_M · 42.8 GB used
⚖ Best $/tok-per-sec
H100 80GB
$30,800
103.8 tok/s · $297 / tok-per-sec

All 3 viable setups — sorted by price

#HardwareQuanttok/sRAM usedPrice$/tok·sWattsConfidence
1RTX 6000 AdaQ4_K_M18.542.8 GB$7,600$410380W○ est.
2A100 80GBQ8_020.478.8 GB$15,800$773480W○ est.
3H100 80GBQ8_0103.878.8 GB$30,800$297780W◐ interp. · vllm

Want the deeper buy-vs-build breakdown?

The AI Hardware Buyer's Guide 2026 covers the hard tradeoffs behind these picks — ROCm vs CUDA real-world cost, used-server bargains, multi-GPU rigs, M5 Max vs RTX workstation, warranty + resale value. Free PDF.

→ Pair this with the LLM Hardware Checker to verify your current machine, or the Self-Host vs API calculator to see whether buying makes sense at all for your workload.

What the LLM hardware recommender does

Pick a model (Llama 3.1, Qwen 2.5, Mistral, DeepSeek), set the minimum tokens/sec you'll tolerate, and this tool returns the cheapest Mac M-series or NVIDIA / AMD GPU that runs it smoothly — backed by 80+ measured benchmarks rather than spec-sheet guesses. It answers "what's the cheapest thing that runs this model fast enough for me?"

How to find your hardware

  1. Choose the model (and size) you want to run.
  2. Set your minimum acceptable speed in tokens per second.
  3. Get the cheapest Mac or GPU that clears that bar, with the measured throughput.

What decides whether a model runs

Two things: VRAM / unified memory — the model's weights (at your quantization) have to fit, or it won't load at all — and memory bandwidth, which sets how fast tokens come out once it fits. That's why an Apple M-series with lots of unified memory can run larger models than a GPU with less VRAM, even if the GPU is "faster" on paper. To decide whether to buy at all, run the self-host vs API calculator; the raw numbers live in the local LLM benchmark database.

Frequently asked questions

How much VRAM do I need for a local LLM?

Roughly the model's parameter count × bytes-per-parameter at your quantization — e.g. an 8B model at 4-bit needs ~5–6 GB. The tool picks hardware that fits your chosen model.

Mac or NVIDIA for local LLMs?

Macs win on memory capacity per dollar (run bigger models); NVIDIA wins on raw throughput. The recommender compares both against your speed target.

Where do the numbers come from?

80+ real measured benchmarks in the site's benchmark database, not vendor spec sheets.

SharePost

More tools like this