#quantization

6 posts tagged with #quantization

Every article below is hand-written, technically reviewed, and focused on quantization. Posts cover real-world architecture decisions, code-level implementation patterns, and trade-offs you'll only discover after shipping production systems.

LM Studio app desktop client screen local LLM — illustration for article on LM Studio Power Developer Tools

LM Studio Power User Setup 2026: Profiles, Quants, LAN Serving

My opinionated 2026 LM Studio setup: reproducible model profiles, quant hygiene rules, scripted loads via lms + /api/v1, and a secure multi-user LAN server.

A computer monitor sitting on top of a desk AI and Machine Learning

How to Run Qwen 35B on 16GB VRAM [2026]: Flags + Quants

A reproducible 16GB recipe for “Qwen 35B”: which Qwen2.5-32B quants fit, how to budget KV cache, and the exact serving flags that stop OOMs.

Computer screen displaying code and terminal prompts Technology

Local LLM Benchmark Methodology [2026]: TTFT vs tok/s Done Right

Stop screenshot-benchmarking. Here’s a reproducible local LLM benchmark methodology for 2026 that separates TTFT from throughput and reports rerunnable results.

a close up of a cpu chip on a table AI and Machine Learning

Gemma 4 26B CPU Inference Benchmark: 5 tok/s Production Math [2026]

A $300 Xeon from 2013 runs Gemma 4 26B at 5 tok/s with no GPU. Here's the memory bandwidth math, quantization tradeoffs, and production decision framework nobody else is covering.

selective focus photography of GEFORCE RTX graphics card AI and Machine Learning

LLM Quantization Levels Compared: Q4_K_M vs Q8_0 vs FP16 [2026]

The practitioner's guide to choosing between Q4_K_M, Q5_K_S, Q8_0, and FP16 quantization for local LLMs — with real perplexity numbers, throughput benchmarks, and per-use-case recommendations.

GGUF vs GPTQ vs EXL2: LLM Quantization Compared [2026] AI and Machine Learning

GGUF vs GPTQ vs EXL2: LLM Quantization Compared [2026]

A head-to-head comparison of GGUF, GPTQ, and EXL2 quantization formats with real quality, speed, and VRAM trade-offs — updated for the 2026 Hugging Face acquisition of ggml.ai.