Self-Host vs API Calculator

Should you buy hardware and run local LLMs, or keep paying API bills? Free calculator: enter your workload, get the cheapest API option vs cheapest viable Mac/GPU setup, break-even month, and 3-year total cost of ownership.

Should you keep paying API bills forever, or buy hardware and run open-weight models locally? Enter your workload below — we'll compute the cheapest API option in your quality tier, the cheapest viable hardware setup, the break-even month, and the 36-month TCO.

Quick presets:

Your workload

Self-host wins
Buying RTX 3090 running Qwen 2.5 Coder 32B saves $3,926 (64%) over 36 months vs paying for Gemini 2.5 Flash. Breaks even in month 11.

API path

$170/month
Using Gemini 2.5 Flash (Google)
$0.3/M in · $2.5/M out
36-month total: $6,120
Other options in this tier
  • Gemini 2.5 Pro (Google) — $688/mo
  • GPT-4.1 (OpenAI) — $700/mo
  • Claude Sonnet 4.6 (Anthropic) — $1,200/mo

Self-host path

$1,600 upfront
RTX 3090 running Qwen 2.5 Coder 32B (Q4_K_M)
39.3 tok/s · ○ estimated
Power: ~103.2 kWh/mo = $17/mo
36-month total: $2,194
Other saturating hardware options
  • RX 7900 XTX$1,650 upfront, 40.3 tok/s
  • RTX 4090$2,600 upfront, 42 tok/s
  • RTX 5090$3,200 upfront, 56.4 tok/s
  • A100 40GB$9,800 upfront, 33.9 tok/s

Cumulative cost over time

$0.000$3,060$6,120Break-even: month 11m0m9m18m27m36
API (Gemini 2.5 Flash) Self-host (RTX 3090) Break-even

The full hardware decision matrix

The AI Hardware Buyer's Guide 2026 covers the decisions behind these numbers: RTX 5090 vs 4090 vs M5 Max, multi-GPU rigs, used-server bargains, ROCm vs CUDA real-world cost, and per-workload benchmarks for chat / RAG / fine-tuning. Free PDF.

How these numbers were calculated

API cost = (input_tokens × input_price + output_tokens × output_price) × queries / 1M. We pick the cheapest API model in your quality tier (see crosswalk below).

Self-host cost = upfront hardware + (power_watts × hours/day × 30 days ÷ 1000) × $/kWh, summed over your horizon. Hardware prices are 2026 street/refurbished estimates; power is sustained inference draw at batch=1. Mac prices include the full system; PC prices include a $800 base build.

Quality tier crosswalk — we map API tiers to open-weight equivalents based on published benchmarks (MMLU, GPQA, HumanEval, MT-Bench). Frontier = 405B / DeepSeek-V3 class. Strong = 70B class. Good = 32B class. Small = 7-9B class. Not a scientific claim — buying-decision-grade.

Throughput saturation check — your monthly queries × output tokens ÷ seconds-per-month gives required tok/s. If the best single machine can't sustain that, we warn you that you'd need multiple machines (and the calculator doesn't model that yet).

What's NOT included: engineering / ops time, downtime, request burstiness, multi-GPU speedup, batching efficiency, model fine-tuning costs, prompt-caching API discounts. For a hobby / indie / single-team workload these are mostly second-order.

See the underlying tokens/sec data at /llm-benchmarks. Pair this calculator with the LLM Hardware Checker to verify your specific hardware can actually run the recommended model.

SharePost

More tools like this